Multimodal representation learning over heterogeneous networks for tag-based music retrieval

被引:0
|
作者
Mendes da Silva, Angelo Cesar [1 ]
Silva, Diego Furtado [2 ]
Marcacini, Ricardo Marcondes [1 ]
机构
[1] Univ Sao Paulo, Inst Math & Comp Sci, Sao Paulo, SP, Brazil
[2] Univ Fed Sao Carlos, Dept Comp, Sao Carlos, SP, Brazil
基金
巴西圣保罗研究基金会;
关键词
Music representation learning; Multimodal representation learning; Music information retrieval; Tag-based music retrieval; FUSION;
D O I
10.1016/j.eswa.2022.117969
中图分类号
TP18 [人工智能理论];
学科分类号
081104 ; 0812 ; 0835 ; 1405 ;
摘要
Learning how to represent data represented by features obtained from multiple modalities through representation learning strategies has received much attention in Music Information Retrieval. Among several sources of information, musical data can be represented mainly by features extracted from acoustic content, lyrics, and metadata that concentrate complementary information and have relevance when discriminating the recordings. In this work, we propose a new method for learning multimodal representations structured as a heterogeneous network capable of incorporating different musical features in constructing a representation and exploring the similarity simultaneously. Our multimodal representation is centered on the information of tags extracted from a state-of-the-art neural language model and, in a complementary way, the audio represented by the melspectrogram. We submitted our method to a robust evaluation process composed of 10,000 queries with different scenarios and model parameter variations. Besides, we compute the Mean Average Precision and compare the representation proposed to representations built only with audio or tags obtained from a pre-trained neural model. The proposed method achieves the best results in all evaluated scenarios and emphasizes the discriminative power of multimodality can add to musical representations.
引用
收藏
页数:9
相关论文
共 50 条
  • [1] Multimodal representation learning over heterogeneous networks for tag-based music retrieval
    da Silva, Angelo Cesar Mendes
    Silva, Diego Furtado
    Marcacini, Ricardo Marcondes
    [J]. Expert Systems with Applications, 2022, 207
  • [2] MULTIMODAL METRIC LEARNING FOR TAG-BASED MUSIC RETRIEVAL
    Won, Minz
    Oramas, Sergio
    Nieto, Oriol
    Gouyon, Fabien
    Serra, Xavier
    [J]. 2021 IEEE INTERNATIONAL CONFERENCE ON ACOUSTICS, SPEECH AND SIGNAL PROCESSING (ICASSP 2021), 2021, : 591 - 595
  • [3] Joint Hypergraph Learning for Tag-Based Image Retrieval
    Wang, Yaxiong
    Zhu, Li
    Qian, Xueming
    Han, Junwei
    [J]. IEEE TRANSACTIONS ON IMAGE PROCESSING, 2018, 27 (09) : 4437 - 4451
  • [4] A survey of tag-based information retrieval
    Lee S.
    Masoud M.
    Balaji J.
    Belkasim S.
    Sunderraman R.
    Moon S.-J.
    [J]. International Journal of Multimedia Information Retrieval, 2017, 6 (2) : 99 - 113
  • [5] Tag-based Personalized Music Recommendation
    Wang, Mengsha
    Xiao, Yingyuan
    Zheng, Wenguang
    Jiao, Xu
    Hsu, Ching-Hsien
    [J]. 2018 15TH INTERNATIONAL SYMPOSIUM ON PERVASIVE SYSTEMS, ALGORITHMS AND NETWORKS (I-SPAN 2018), 2018, : 201 - 208
  • [6] Large-scale Tag-based Font Retrieval with Generative Feature Learning
    Chen, Tianlang
    Wang, Zhaowen
    Xu, Ning
    Jin, Hailin
    Luo, Jiebo
    [J]. 2019 IEEE/CVF INTERNATIONAL CONFERENCE ON COMPUTER VISION (ICCV 2019), 2019, : 9115 - 9124
  • [7] Adaptive diversification for tag-based social image retrieval
    Ksibi, Amel
    Ben Ammar, Anis
    Ben Amar, Chokri
    [J]. INTERNATIONAL JOURNAL OF MULTIMEDIA INFORMATION RETRIEVAL, 2014, 3 (01) : 29 - 39
  • [8] Tag-Based Social Image Retrieval: An Empirical Evaluation
    Sun, Aixin
    Bhowmick, Sourav S.
    Khanh Tran Nam Nguyen
    Bai, Ge
    [J]. JOURNAL OF THE AMERICAN SOCIETY FOR INFORMATION SCIENCE AND TECHNOLOGY, 2011, 62 (12): : 2364 - 2381
  • [9] A document expansion framework for tag-based image retrieval
    Lu, Wei
    Ding, Heng
    Jiang, Jiepu
    [J]. ASLIB JOURNAL OF INFORMATION MANAGEMENT, 2018, 70 (01) : 47 - 65
  • [10] A Tag-Based Approach for Learning Ergonomic Concepts
    Tsai, Li Chen
    Tang, Kuo Hao
    Hwang, Sheue Ling
    [J]. HUMAN FACTORS AND ERGONOMICS IN MANUFACTURING & SERVICE INDUSTRIES, 2014, 24 (05) : 574 - 584