MKDAT: Multi-Level Knowledge Distillation with Adaptive Temperature for Distantly Supervised Relation Extraction

被引：0

作者：

Long, Jun ^{[1
]}

Yin, Zhuoying ^{[1
,2
]}

Han, Yan ^{[3
]}

Huang, Wenti ^{[4
]}

机构：

[1] Cent South Univ, Big Data Inst, Changsha 410075, Peoples R China

[2] Guizhou Rural Credit Union, Guiyang 550000, Peoples R China

[3] Guizhou Univ Commerce, Sch Comp & Informat Engn, Guiyang 550025, Peoples R China

[4] Hunan Univ Sci & Technol, Sch Comp Sci & Engn, Xiangtan 411100, Peoples R China

来源：

INFORMATION | 2024年 / 15卷 / 07期

基金：

中国国家自然科学基金;

关键词：

distantly supervised relation extraction; knowledge distillation; label softening;

D O I：

10.3390/info15070382

中图分类号：

TP [自动化技术、计算机技术];

学科分类号：

0812 ;

摘要：

Distantly supervised relation extraction (DSRE), first used to address the limitations of manually annotated data via automatically annotating the data with triplet facts, is prone to issues such as mislabeled annotations due to the interference of noisy annotations. To address the interference of noisy annotations, we leveraged a novel knowledge distillation (KD) method which was different from the conventional models on DSRE. More specifically, we proposed a model-agnostic KD method, Multi-Level Knowledge Distillation with Adaptive Temperature (MKDAT), which mainly involves two modules: Adaptive Temperature Regulation (ATR) and Multi-Level Knowledge Distilling (MKD). ATR allocates adaptive entropy-based distillation temperatures to different training instances for providing a moderate softening supervision to the student, in which label hardening is possible for instances with great entropy. MKD combines the bag-level and instance-level knowledge of the teacher as supervisions of the student, and trains the teacher and student at the bag and instance levels, respectively, which aims at mitigating the effects of noisy annotation and improving the sentence-level prediction performance. In addition, we implemented three MKDAT models based on the CNN, PCNN, and ATT-BiLSTM neural networks, respectively, and the experimental results show that our distillation models outperform the baseline models on bag-level and instance-level evaluations.

引用

页数：18

共 50 条

[1] A Multi-teacher Knowledge Distillation Framework for Distantly Supervised Relation Extraction with Flexible Temperature
Fei, Hongxiao
Tan, Yangying
Huang, Wenti
Long, Jun
Huang, Jincai
Yang, Liu
WEB AND BIG DATA, PT II, APWEB-WAIM 2023, 2024, 14332 : 103 - 116
[2] Improving Distantly Supervised Relation Extraction with Multi-Level Noise Reduction
Song, Wei
Yang, Zijiang
AI, 2024, 5 (03) : 1709 - 1730
[3] Distantly Supervised relation extraction with multi-level contextual information integration
Han, Danjie
Huang, Heyan
Shi, Shumin
Yuan, Changsen
Guo, Cunhan
NEUROCOMPUTING, 2025, 634
[4] Multi-Level Structured Self-Attentions for Distantly Supervised Relation Extraction
Du, Jinhua
Han, Jingguang
Way, Andy
Wan, Dadong
2018 CONFERENCE ON EMPIRICAL METHODS IN NATURAL LANGUAGE PROCESSING (EMNLP 2018), 2018, : 2216 - 2225
[5] Adaptive multi-teacher multi-level knowledge distillation
Liu, Yuang
Zhang, Wei
Wang, Jun
NEUROCOMPUTING, 2020, 415 : 106 - 113
[6] Adaptive multi-teacher multi-level knowledge distillation
Liu, Yuang
Zhang, Wei
Wang, Jun
Neurocomputing, 2021, 415 : 106 - 113
[7] MiDTD: A Simple and Effective Distillation Framework for Distantly Supervised Relation Extraction
Li, Rui
Yang, Cheng
Li, Tingwei
Su, Sen
ACM TRANSACTIONS ON INFORMATION SYSTEMS, 2022, 40 (04)
[8] Distantly supervised Web relation extraction for knowledge base population
Augenstein, Isabelle
Maynard, Diana
Ciravegna, Fabio
SEMANTIC WEB, 2016, 7 (04) : 335 - 349
[9] Hierarchical Knowledge Transfer Network for Distantly Supervised Relation Extraction
Song, Wei
Gu, Weishuai
2023 INTERNATIONAL JOINT CONFERENCE ON NEURAL NETWORKS, IJCNN, 2023,
[10] Knowledge-embodied attention for distantly supervised relation extraction
Deng, Kejun
Zhang, Xuemiao
Ye, Songtao
Liu, Junfei
INTELLIGENT DATA ANALYSIS, 2020, 24 (02) : 445 - 457

← 1 2 3 4 5 →