Multiple Bayesian discriminant functions for high-dimensional massive data classification

被引:0
|
作者
Jianfei Zhang
Shengrui Wang
Lifei Chen
Patrick Gallinari
机构
[1] Université de Sherbrooke,ProspectUS Laboratoire, Département d’Informatique
[2] Fujian Normal University,School of Mathematics and Computer Science
[3] Université Pierre et Marie Curie,Laboratoire d’Informatique de Paris 6 (LIP6)
来源
关键词
Decision boundaries; Naive Bayes; Feature weighting; High-dimensional massive data; Class dispersion;
D O I
暂无
中图分类号
学科分类号
摘要
The presence of complex distributions of samples concealed in high-dimensional, massive sample-size data challenges all of the current classification methods for data mining. Samples within a class usually do not uniformly fill a certain (sub)space but are individually concentrated in certain regions of diverse feature subspaces, revealing the class dispersion. Current classifiers applied to such complex data inherently suffer from either high complexity or weak classification ability, due to the imbalance between flexibility and generalization ability of the discriminant functions used by these classifiers. To address this concern, we propose a novel representation of discriminant functions in Bayesian inference, which allows multiple Bayesian decision boundaries per class, each in its individual subspace. For this purpose, we design a learning algorithm that incorporates the naive Bayes and feature weighting approaches into structural risk minimization to learn multiple Bayesian discriminant functions for each class, thus combining the simplicity and effectiveness of naive Bayes and the benefits of feature weighting in handling high-dimensional data. The proposed learning scheme affords a recursive algorithm for exploring class density distribution for Bayesian estimation, and an automated approach for selecting powerful discriminant functions while keeping the complexity of the classifier low. Experimental results on real-world data characterized by millions of samples and features demonstrate the promising performance of our approach.
引用
收藏
页码:465 / 501
页数:36
相关论文
共 50 条
  • [1] Multiple Bayesian discriminant functions for high-dimensional massive data classification
    Zhang, Jianfei
    Wang, Shengrui
    Chen, Lifei
    Gallinari, Patrick
    DATA MINING AND KNOWLEDGE DISCOVERY, 2017, 31 (02) : 465 - 501
  • [2] On sparse linear discriminant analysis algorithm for high-dimensional data classification
    Ng, Michael K.
    Liao, Li-Zhi
    Zhang, Leihong
    NUMERICAL LINEAR ALGEBRA WITH APPLICATIONS, 2011, 18 (02) : 223 - 235
  • [3] Modified linear discriminant analysis approaches for classification of high-dimensional microarray data
    Xu, Ping
    Brock, Guy N.
    Parrish, Rudolph S.
    COMPUTATIONAL STATISTICS & DATA ANALYSIS, 2009, 53 (05) : 1674 - 1687
  • [4] QUADRATIC DISCRIMINANT ANALYSIS FOR HIGH-DIMENSIONAL DATA
    Wu, Yilei
    Qin, Yingli
    Zhu, Mu
    STATISTICA SINICA, 2019, 29 (02) : 939 - 960
  • [5] Bayesian weighted random forest for classification of high-dimensional genomics data
    Olaniran, Oyebayo Ridwan
    Abdullah, Mohd Asrul A.
    KUWAIT JOURNAL OF SCIENCE, 2023, 50 (04) : 477 - 484
  • [6] A Hybrid Dimension Reduction Based Linear Discriminant Analysis for Classification of High-Dimensional Data
    Zorarpaci, Ezgi
    2021 IEEE CONGRESS ON EVOLUTIONARY COMPUTATION (CEC 2021), 2021, : 1028 - 1036
  • [7] CLASSIFICATION OF HIGH-DIMENSIONAL DATA: A RANDOM-MATRIX REGULARIZED DISCRIMINANT ANALYSIS APPROACH
    Ye, Bin
    Liu, Peng
    INTERNATIONAL JOURNAL OF INNOVATIVE COMPUTING INFORMATION AND CONTROL, 2019, 15 (03): : 955 - 967
  • [8] Bayesian clinical classification from high-dimensional data: Signatures versus variability
    Shalabi, Akram
    Inoue, Masato
    Watkins, Johnathan
    De Rinaldis, Emanuele
    Coolen, Anthony C. C.
    STATISTICAL METHODS IN MEDICAL RESEARCH, 2018, 27 (02) : 336 - 351
  • [9] MULTICATEGORY VERTEX DISCRIMINANT ANALYSIS FOR HIGH-DIMENSIONAL DATA
    Wu, Tong Tong
    Lange, Kenneth
    ANNALS OF APPLIED STATISTICS, 2010, 4 (04): : 1698 - 1721
  • [10] A modified linear discriminant analysis for high-dimensional data
    Hyodo, Masashi
    Yamada, Takayuki
    Himeno, Tetsuto
    Seo, Takashi
    HIROSHIMA MATHEMATICAL JOURNAL, 2012, 42 (02) : 209 - 231