MonoATT: Online Monocular 3D Object Detection with Adaptive Token Transformer

被引：11

作者：

Zhou, Yunsong ^{[1
]}

Zhu, Hongzi ^{[1
]}

Liu, Quan ^{[1
]}

Chang, Shan ^{[2
]}

Guo, Minyi ^{[1
]}

机构：

[1] Shanghai Jiao Tong Univ, Shanghai, Peoples R China

[2] Donghua Univ, Shanghai, Peoples R China

来源：

2023 IEEE/CVF CONFERENCE ON COMPUTER VISION AND PATTERN RECOGNITION (CVPR) | 2023年

基金：

上海市自然科学基金; 中国国家自然科学基金;

关键词：

D O I：

10.1109/CVPR52729.2023.01678

中图分类号：

TP18 [人工智能理论];

学科分类号：

081104 ; 0812 ; 0835 ; 1405 ;

摘要：

Mobile monocular 3D object detection (Mono3D) (e.g., on a vehicle, a drone, or a robot) is an important yet challenging task. Existing transformer-based offline Mono3D models adopt grid-based vision tokens, which is suboptimal when using coarse tokens due to the limited available computational power. In this paper, we propose an online Mono3D framework, called MonoATT, which leverages a novel vision transformer with heterogeneous tokens of varying shapes and sizes to facilitate mobile Mono3D. The core idea of MonoATT is to adaptively assign finer tokens to areas of more significance before utilizing a transformer to enhance Mono3D. To this end, we first use prior knowledge to design a scoring network for selecting the most important areas of the image, and then propose a token clustering and merging network with an attention mechanism to gradually merge tokens around the selected areas in multiple stages. Finally, a pixel-level feature map is reconstructed from heterogeneous tokens before employing a SOTA Mono3D detector as the underlying detection core. Experiment results on the real-world KITTI dataset demonstrate that MonoATT can effectively improve the Mono3D accuracy for both near and far objects and guarantee low latency. MonoATT yields the best performance compared with the state-of-the-art methods by a large margin and is ranked number one on the KITTI 3D benchmark.

引用

页码：17493 / 17503

页数：11

共 50 条

[21] Progressive Coordinate Transforms for Monocular 3D Object Detection
Wang, Li
Zhang, Li
Zhu, Yi
Zhang, Zhi
He, Tong
Li, Mu
Xue, Xiangyang
ADVANCES IN NEURAL INFORMATION PROCESSING SYSTEMS 34 (NEURIPS 2021), 2021, 34
[22] Exploring Geometric Consistency for Monocular 3D Object Detection
Lian, Qing
Ye, Botao
Xu, Ruijia
Yao, Weilong
Zhang, Tong
2022 IEEE/CVF CONFERENCE ON COMPUTER VISION AND PATTERN RECOGNITION (CVPR 2022), 2022, : 1675 - 1684
[23] MonoSG: Monocular 3D Object Detection With Stereo Guidance
Fan, Zhiwei
Xu, Chao
Chu, Minghang
Huang, Yuling
Ma, Yaoyao
Wang, Jing
Xu, Yishen
Wu, Di
IEEE ROBOTICS AND AUTOMATION LETTERS, 2025, 10 (04): : 3604 - 3611
[24] Monocular 3D Object Detection With Motion Feature Distillation
Hu, Henan
Li, Muyu
Zhu, Ming
Gao, Wen
Liu, Peiyu
Chan, Kwok-Leung
IEEE ACCESS, 2023, 11 : 82933 - 82945
[25] Monocular Object Detection Using 3D Geometric Primitives
Carr, Peter
Sheikh, Yaser
Matthews, Iain
COMPUTER VISION - ECCV 2012, PT I, 2012, 7572 : 864 - 878
[26] Monocular 3D Object Detection from Roadside Infrastructure
Huang, Delu
Wen, Feng
2024 35TH IEEE INTELLIGENT VEHICLES SYMPOSIUM, IEEE IV 2024, 2024, : 1672 - 1677
[27] Dense-JANet for Monocular 3D Object Detection
Shang, Xiaoqing
Cheng, Zhiwei
Shi, Su
Cheng, Zhuanghao
Huang, Hongcheng
2020 IEEE 23RD INTERNATIONAL CONFERENCE ON INTELLIGENT TRANSPORTATION SYSTEMS (ITSC), 2020,
[28] DST3D: DLA-Swin Transformer for Single-Stage Monocular 3D Object Detection
Wu, Zhihong
Jiang, Xin
Xu, Ruidong
Lu, Ke
Zhu, Yuan
Wu, Mingzhi
2022 IEEE INTELLIGENT VEHICLES SYMPOSIUM (IV), 2022, : 411 - 418
[29] Monocular 3D object detection for an indoor robot environment
Kim, Jiwon
Lee, GiJae
Kim, Jun-Sik
Kim, Hyunwoo J.
Kim, KangGeon
2020 29TH IEEE INTERNATIONAL CONFERENCE ON ROBOT AND HUMAN INTERACTIVE COMMUNICATION (RO-MAN), 2020, : 438 - 445
[30] MonoCD: Monocular 3D Object Detection with Complementary Depths
Yan, Longfei
Yan, Pei
Xiong, Shengzhou
Xiang, Xuanyu
Tan, Yihua
2024 IEEE/CVF CONFERENCE ON COMPUTER VISION AND PATTERN RECOGNITION (CVPR), 2024, : 10248 - 10257

← 1 2 3 4 5 →