Feature selection using Fisher score and multilabel neighborhood rough sets for multilabel classification

被引：143

作者：

Sun, Lin ^{[1
,3
,4
]}

Wang, Tianxiang ^{[1
]}

Ding, Weiping ^{[2
]}

Xu, Jiucheng ^{[1
,4
]}

Lin, Yaojin ^{[3
]}

机构：

[1] Henan Normal Univ, Coll Comp & Informat Engn, Xinxiang 453007, Henan, Peoples R China

[2] Nantong Univ, Sch Informat Sci & Technol, Nantong 226019, Peoples R China

[3] Minnan Normal Univ, Key Lab Data Sci & Intelligence Applicat, Zhangzhou 363000, Peoples R China

[4] Key Lab Artificial Intelligence & Personalized Le, Xinxiang 453007, Henan, Peoples R China

来源：

INFORMATION SCIENCES | 2021年 / 578卷

基金：

中国国家自然科学基金;

关键词：

Feature selection; Neighborhood rough sets; Fisher Score; Multilabel classification; LABEL FEATURE-SELECTION; UNCERTAINTY MEASURES; INFORMATION;

D O I：

10.1016/j.ins.2021.08.032

中图分类号：

TP [自动化技术、计算机技术];

学科分类号：

0812 ;

摘要：

In recent years, feature selection for multilabel classification has attracted attention in machine learning and data mining. However, some feature selection methods ignore the correlations among labels, resulting in low performance, and most of them face challenges in determining an appropriate neighborhood radius for neighborhood systems and suffer from expensive time cost. To overcome the issues, we propose a novel feature selection method using Fisher score and multilabel neighborhood rough sets (MNRS) in multilabel neighborhood decision systems. First, to identify the correlations between labels under a binary distribution, two types of new mutual information between labels are considered, and their balance coefficients are defined. By enhancing strong correlations and weakening weak correlations between labels, a mutual information-based Fisher score model with a second-order correlation between labels is designed to fit multilabel data. Second, to address the problem of automatically choosing a neighborhood radius, a subset of hetero-geneous and homogeneous samples is employed to develop a new classification margin as a neighborhood radius, and some concepts of neighborhood, neighborhood class, and upper and lower approximations are formulated for multilabel neighborhood decision systems. The weight and dependency degree are presented to effectively measure the uncertainty of samples in multilabel neighborhood decision systems. Thus, we further present a new classification margin-based MNRS model. Finally, a filter-wrapper preprocessing algorithm for feature selection using the improved Fisher score model is proposed to decrease the spatiotemporal complexity of multilabel data, and a heuristic feature selection algorithm is designed for improve classification performance on multilabel datasets. Experimental results on thirteen multilabel datasets show that the proposed algorithm is effective in selecting significant features, demonstrating its excellent classification ability in multilabel datasets. (c) 2021 Elsevier Inc. All rights reserved.

引用

页码：887 / 912

页数：26

共 50 条

[41] Effective Evolutionary Multilabel Feature Selection under a Budget Constraint
Lee, Jaesung
Seo, Wangduk
Kim, Dae-Won
COMPLEXITY, 2018,
[42] Feature selection based on neighborhood rough sets and Gini index
Zhang Y.
Nie B.
Du J.
Chen J.
Du Y.
Jin H.
Zheng X.
Chen X.
Miao Z.
PeerJ Computer Science, 2023, 9
[43] Feature selection based on neighborhood rough sets and Gini index
Zhang, Yuchao
Nie, Bin
Du, Jianqiang
Chen, Jiandong
Du, Yuwen
Jin, Haike
Zheng, Xuepeng
Chen, Xingxin
Miao, Zhen
PEERJ COMPUTER SCIENCE, 2023, 9
[44] Feature selection based on neighborhood rough sets and Gini index
Zhang, Yuchao
Nie, Bin
Du, Jianqiang
Chen, Jiandong
Du, Yuwen
Jin, Haike
Zheng, Xuepeng
Chen, Xingxin
Miao, Zhen
PEERJ, 2023, 11
[45] Neighborhood rough sets with distance metric learning for feature selection
Yang, Xiaoling
Chen, Hongmei
Li, Tianrui
Wan, Jihong
Sang, Binbin
KNOWLEDGE-BASED SYSTEMS, 2021, 224
[46] Exploiting Multilabel Information for Noise-Resilient Feature Selection
Jian, Ling
Li, Jundong
Liu, Huan
ACM TRANSACTIONS ON INTELLIGENT SYSTEMS AND TECHNOLOGY, 2018, 9 (05)
[47] Interactive streaming feature selection based on neighborhood rough sets
Zhang, Gangqiang
Hu, Jingjing
Yang, Jing
Zhang, Pengfei
ENGINEERING APPLICATIONS OF ARTIFICIAL INTELLIGENCE, 2025, 139
[48] Feature subset selection based on fuzzy neighborhood rough sets
Wang, Changzhong
Shao, Mingwen
He, Qiang
Qian, Yuhua
Qi, Yali
KNOWLEDGE-BASED SYSTEMS, 2016, 111 : 173 - 179
[49] Feature selection for imbalanced data based on neighborhood rough sets
Chen, Hongmei
Li, Tianrui
Fan, Xin
Luo, Chuan
INFORMATION SCIENCES, 2019, 483 : 1 - 20
[50] Multilabel Text Classification Using Multilayer DGAT
Chen, Hui
Huang, Jian
Tao, Nana
Huang, Jijie
Wang, Jing
IEEE ACCESS, 2022, 10 : 125042 - 125051

← 1 2 3 4 5 →