Multiple Bayesian discriminant functions for high-dimensional massive data classification

被引：0

作者：

Jianfei Zhang

Shengrui Wang

Lifei Chen

Patrick Gallinari

机构：

[1] Université de Sherbrooke,ProspectUS Laboratoire, Département d’Informatique

[2] Fujian Normal University,School of Mathematics and Computer Science

[3] Université Pierre et Marie Curie,Laboratoire d’Informatique de Paris 6 (LIP6)

来源：

Data Mining and Knowledge Discovery | 2017年 / 31卷

关键词：

Decision boundaries; Naive Bayes; Feature weighting; High-dimensional massive data; Class dispersion;

D O I：

暂无

中图分类号：

学科分类号：

摘要：

The presence of complex distributions of samples concealed in high-dimensional, massive sample-size data challenges all of the current classification methods for data mining. Samples within a class usually do not uniformly fill a certain (sub)space but are individually concentrated in certain regions of diverse feature subspaces, revealing the class dispersion. Current classifiers applied to such complex data inherently suffer from either high complexity or weak classification ability, due to the imbalance between flexibility and generalization ability of the discriminant functions used by these classifiers. To address this concern, we propose a novel representation of discriminant functions in Bayesian inference, which allows multiple Bayesian decision boundaries per class, each in its individual subspace. For this purpose, we design a learning algorithm that incorporates the naive Bayes and feature weighting approaches into structural risk minimization to learn multiple Bayesian discriminant functions for each class, thus combining the simplicity and effectiveness of naive Bayes and the benefits of feature weighting in handling high-dimensional data. The proposed learning scheme affords a recursive algorithm for exploring class density distribution for Bayesian estimation, and an automated approach for selecting powerful discriminant functions while keeping the complexity of the classifier low. Experimental results on real-world data characterized by millions of samples and features demonstrate the promising performance of our approach.

引用

页码：465 / 501

页数：36

共 50 条

[41] Adaptive Bayesian density regression for high-dimensional data
Shen, Weining
Ghosal, Subhashis
BERNOULLI, 2016, 22 (01) : 396 - 420
[42] Sparse Bayesian multinomial probit regression model with correlation prior for high-dimensional data classification
Yang Aijun
Jiang Xuejun
Liu Pengfei
Lin Jinguan
STATISTICS & PROBABILITY LETTERS, 2016, 119 : 241 - 247
[43] Multiple imputation in the presence of high-dimensional data
Zhao, Yize
Long, Qi
STATISTICAL METHODS IN MEDICAL RESEARCH, 2016, 25 (05) : 2021 - 2035
[44] Multiple imputation with compatibility for high-dimensional data
Zahid, Faisal Maqbool
Faisal, Shahla
Heumann, Christian
PLOS ONE, 2021, 16 (07):
[45] Class-specific subspace discriminant analysis for high-dimensional data
Bouveyron, Charles
Girard, Stephane
Schmid, Cordelia
SUBSPACE, LATENT STRUCTURE AND FEATURE SELECTION, 2006, 3940 : 139 - 150
[46] Graph-based sparse linear discriminant analysis for high-dimensional classification
Liu, Jianyu
Yu, Guan
Liu, Yufeng
JOURNAL OF MULTIVARIATE ANALYSIS, 2019, 171 : 250 - 269
[47] High-dimensional image data feature extraction by double discriminant embedding
Maryam Imani
Hassan Ghassemian
Pattern Analysis and Applications, 2017, 20 : 473 - 484
[48] High-dimensional image data feature extraction by double discriminant embedding
Imani, Maryam
Ghassemian, Hassan
PATTERN ANALYSIS AND APPLICATIONS, 2017, 20 (02) : 473 - 484
[49] Data-dependent kernels for high-dimensional data classification
Wang, JD
Kwok, JT
Shen, HC
Quan, L
PROCEEDINGS OF THE INTERNATIONAL JOINT CONFERENCE ON NEURAL NETWORKS (IJCNN), VOLS 1-5, 2005, : 102 - 107
[50] High-Dimensional Bayesian Geostatistics
Banerjee, Sudipto
BAYESIAN ANALYSIS, 2017, 12 (02): : 583 - 614

← 1 2 3 4 5 →