Listen to Look: Action Recognition by Previewing Audio

被引:132
|
作者
Gao, Ruohan [1 ,2 ]
Oh, Tae-Hyun [2 ,3 ]
Grauman, Kristen [1 ,2 ]
Torresani, Lorenzo [2 ]
机构
[1] Univ Texas Austin, Austin, TX 78712 USA
[2] Facebook AI Res, Austin, TX 78701 USA
[3] POSTECH, Dept EE, Pohang, South Korea
关键词
D O I
10.1109/CVPR42600.2020.01047
中图分类号
TP18 [人工智能理论];
学科分类号
081104 ; 0812 ; 0835 ; 1405 ;
摘要
In the face of the video data deluge, today's expensive clip-level classifiers are increasingly impractical. We propose a framework for efficient action recognition in untrimmed video that uses audio as a preview mechanism to eliminate both short-term and long-term visual redun-dancies. First, we devise an IMGAUD2VID framework that hallucinates clip-level features by distilling from lighter modalities-a single frame and its accompanying audio-reducing short-term temporal redundancy for efficient clip-level recognition. Second, building on IMGAUD2VID, we further propose IMGAUD-SKIMMING, an attention-based long short-term memory network that iteratively selects useful moments in untrimmed videos, reducing long-term temporal redundancy for efficient video-level recognition. Extensive experiments on four action recognition datasets demonstrate that our method achieves the state-of-the-art in terms of both recognition accuracy and speed.
引用
收藏
页码:10454 / 10464
页数:11
相关论文
共 50 条
  • [41] Listen with Intent: Improving Speech Recognition with Audio-to-Intent Front-End
    Ray, Swayambhu Nath
    Wu, Minhua
    Raju, Anirudh
    Ghahremani, Pegah
    Bilgi, Raghavendra
    Rao, Milind
    Arsikere, Harish
    Rastrow, Ariya
    Stolcke, Andreas
    Droppo, Jasha
    INTERSPEECH 2021, 2021, : 3455 - 3459
  • [42] Effects of Initial Testing and Previewing on Recognition
    Belchev, Zorry
    Bodner, Glen E.
    CANADIAN JOURNAL OF EXPERIMENTAL PSYCHOLOGY-REVUE CANADIENNE DE PSYCHOLOGIE EXPERIMENTALE, 2014, 68 (04): : 294 - 294
  • [43] A Closer Look at Spatiotemporal Convolutions for Action Recognition
    Tran, Du
    Wang, Heng
    Torresani, Lorenzo
    Ray, Jamie
    LeCun, Yann
    Paluri, Manohar
    2018 IEEE/CVF CONFERENCE ON COMPUTER VISION AND PATTERN RECOGNITION (CVPR), 2018, : 6450 - 6459
  • [44] STOP, LOOK, AND LISTEN - COMMENTARY
    BARAN, N
    BYTE, 1995, 20 (08): : 218 - 218
  • [45] Look at the Valve but Listen to the Patient
    Patel, Himanshu J.
    CIRCULATION-CARDIOVASCULAR QUALITY AND OUTCOMES, 2019, 12 (02):
  • [46] DONT LOOK IT UP - LISTEN
    WEAVER, CH
    SPEECH TEACHER, 1957, 6 (03): : 240 - 246
  • [47] Freud, from look to listen
    Delon, Michel
    EUROPE-REVUE LITTERAIRE MENSUELLE, 2019, (1082) : 371 - 372
  • [48] Look, listen, and learn to read
    Collins, NLD
    Shaeffer, MB
    YOUNG CHILDREN, 1997, 52 (05): : 65 - 68
  • [49] Re: Listen, look, think
    Jorgensen, Arne C.
    TIDSSKRIFT FOR DEN NORSKE LAEGEFORENING, 2015, 135 (17) : 1534 - 1534
  • [50] A Closer Look at Video Sampling for Sequential Action Recognition
    Zhang, Yu
    Zhao, Junjie
    Chen, Zhengjie
    Mi, Siya
    Zhu, Hongyuan
    Geng, Xin
    IEEE TRANSACTIONS ON CIRCUITS AND SYSTEMS FOR VIDEO TECHNOLOGY, 2023, 33 (12) : 7503 - 7514