Collaborative Noisy Label Cleaner: Learning Scene-aware Trailers for Multi-modal Highlight Detection in Movies

被引:4
|
作者
Gan, Bei [1 ]
Shu, Xiujun [1 ]
Qiao, Ruizhi [1 ]
Wu, Haoqian [1 ]
Chen, Keyu [1 ]
Li, Hanjun [1 ]
Ren, Bo [1 ]
机构
[1] Tencent YouTu Lab, Shanghai, Peoples R China
关键词
D O I
10.1109/CVPR52729.2023.01812
中图分类号
TP18 [人工智能理论];
学科分类号
081104 ; 0812 ; 0835 ; 1405 ;
摘要
Movie highlights stand out of the screenplay for efficient browsing and play a crucial role on social media platforms. Based on existing efforts, this work has two observations: (1) For different annotators, labeling highlight has uncertainty, which leads to inaccurate and time-consuming annotations. (2) Besides previous supervised or unsupervised settings, some existing video corpora can be useful, e.g., trailers, but they are often noisy and incomplete to cover the full highlights. In this work, we study a more practical and promising setting, i.e., reformulating highlight detection as "learning with noisy labels". This setting does not require time-consuming manual annotations and can fully utilize existing abundant video corpora. First, based on movie trailers, we leverage scene segmentation to obtain complete shots, which are regarded as noisy labels. Then, we propose a Collaborative noisy Label Cleaner (CLC) framework to learn from noisy highlight moments. CLC consists of two modules: augmented cross-propagation (ACP) and multi-modality cleaning (MMC). The former aims to exploit the closely related audio-visual signals and fuse them to learn unified multi-modal representations. The latter aims to achieve cleaner highlight labels by observing the changes in losses among different modalities. To verify the effectiveness of CLC, we further collect a large-scale highlight dataset named MovieLights. Comprehensive experiments on MovieLights and YouTube Highlights datasets demonstrate the effectiveness of our approach. Code has been made available at: https://github.com/TencentYoutuResearch/HighlightDetection-CLC.
引用
收藏
页码:18898 / 18907
页数:10
相关论文
共 10 条
  • [1] Scene-Aware Prompt for Multi-modal Dialogue Understanding and Generation
    Li, Bin
    Weng, Yixuan
    Ma, Ziyu
    Sun, Bin
    Li, Shutao
    NATURAL LANGUAGE PROCESSING AND CHINESE COMPUTING, NLPCC 2022, PT II, 2022, 13552 : 179 - 191
  • [2] Scene-Aware Label Graph Learning for Multi-Label Image Classification
    Zhu, Xuelin
    Liu, Jian
    Liu, Weijia
    Ge, Jiawei
    Liu, Bo
    Cao, Jiuxin
    2023 IEEE/CVF INTERNATIONAL CONFERENCE ON COMPUTER VISION, ICCV, 2023, : 1473 - 1482
  • [3] Visual Scene-Aware Hybrid and Multi-Modal Feature Aggregation for Facial Expression Recognition
    Lee, Min Kyu
    Kim, Dae Ha
    Song, Byung Cheol
    SENSORS, 2020, 20 (18) : 1 - 24
  • [4] MMDL: a multi-modal deep learning for video highlight detection in sports
    Qiaoyun Zhang
    Chih-Yung Chang
    Shih-Jung Wu
    Hsiang-Chuan Chang
    Diptendu Sinha Roy
    International Journal of Multimedia Information Retrieval, 2025, 14 (2)
  • [5] TaoHighlight: Commodity-Aware Multi-Modal Video Highlight Detection in E-Commerce
    Guo, Zhaoyu
    Zhao, Zhou
    Jin, Weike
    Wang, Dazhou
    Liu, Ruitao
    Yu, Jun
    IEEE TRANSACTIONS ON MULTIMEDIA, 2022, 24 : 2606 - 2616
  • [6] CONTEXT-AWARE DEEP LEARNING FOR MULTI-MODAL DEPRESSION DETECTION
    Lam, Genevieve
    Huang Dongyan
    Lin, Weisi
    2019 IEEE INTERNATIONAL CONFERENCE ON ACOUSTICS, SPEECH AND SIGNAL PROCESSING (ICASSP), 2019, : 3946 - 3950
  • [7] Scene-adaptive and Region-aware Multi-modal Prompt for Open Vocabulary Object Detection
    Zhao, Xiaowei
    Liu, Xianglong
    Wang, Duorui
    Gao, Yajun
    Liu, Zhide
    2024 IEEE/CVF CONFERENCE ON COMPUTER VISION AND PATTERN RECOGNITION (CVPR), 2024, : 16741 - 16750
  • [8] Joint learning of video scene detection and annotation via multi-modal adaptive context network
    Xu, Yifei
    Pan, Litong
    Sang, Weiguang
    Luo, Hailun
    Li, Li
    Wei, Pingping
    Zhu, Li
    EXPERT SYSTEMS WITH APPLICATIONS, 2024, 249
  • [9] OCR-Aware Scene Graph Generation Via Multi-modal Object Representation Enhancement and Logical Bias Learning
    Zhou, Xinyu
    Ji, Zihan
    Zhu, Anna
    PATTERN RECOGNITION AND COMPUTER VISION, PRCV 2024, PT VII, 2025, 15037 : 201 - 215
  • [10] Context-Aware Inductive Bias Learning for Vessel Border Detection in Multi-modal Intracoronary Imaging
    Gao, Zhifan
    Li, Shuo
    MEDICAL IMAGE COMPUTING AND COMPUTER ASSISTED INTERVENTION - MICCAI 2019, PT II, 2019, 11765 : 776 - 784