Spam Detection Based on Feature Evolution to Deal with Concept Drift

被引:5
|
作者
Henke, Marcia [1 ]
Santos, Eulanda [2 ]
Souto, Eduardo [2 ]
Santin, Altair O. [3 ]
机构
[1] Fed Univ Santa Maria UFSM, Ind Tech Coll Santa Maria CTISM, Santa Maria, RS, Brazil
[2] Fed Univ Amazonas UFAM, Comp Inst ICOMP, Manaus, Amazonas, Brazil
[3] Pontificia Univ Catolica Parana, Curitiba, PR, Brazil
关键词
Computer Security Network; Machine Learning; Concept Drift;
D O I
10.3897/jucs.66284
中图分类号
TP31 [计算机软件];
学科分类号
081202 ; 0835 ;
摘要
Electronic messages are still considered the most significant tools in business and personal applications due to their low cost and easy access. However, e-mails have become a major problem owing to the high amount of junk mail, named spam, which fill the e-mail boxes of users. Several approaches have been proposed to detect spam, such as filters implemented in e-mail servers and user-based spam message classification mechanisms. A major problem with these approaches is spam detection in the presence of concept drift, especially as a result of changes in features over time. To overcome this problem, this work proposes a new spam detection system based on analyzing the evolution of features. The proposed method is divided into three steps: 1) spam classification model training; 2) concept drift detection; and 3) knowledge transfer learning. The first step generates classification models, as commonly conducted in machine learning. The second step introduces a new strategy to avoid concept drift: SFS (Similarity-based Features Selection) that analyzes the evolution of the features taking into account similarity obtained between the feature vectors extracted from training data and test data. Finally, the third step focuses on the following questions: what, how, and when to transfer acquired knowledge? The proposed method is evaluated using two public datasets. The results of the experiments show that it is possible to infer a threshold to detect changes (drift) in order to ensure that the spam classification model is updated through knowledge transfer. Moreover, our anomaly detection system is able to perform spam classification and concept drift detection as two parallel and independent tasks.
引用
收藏
页码:364 / 386
页数:23
相关论文
共 50 条
  • [1] Analysis of the Evolution of Features in Classification Problems with Concept Drift: Application to Spam Detection
    Henke, Marcia
    Souto, Eduardo
    dos Santos, Eulanda M.
    PROCEEDINGS OF THE 2015 IFIP/IEEE INTERNATIONAL SYMPOSIUM ON INTEGRATED NETWORK MANAGEMENT (IM), 2015, : 874 - 877
  • [2] Content-based concept drift detection for Email spam filtering
    Zi Hayat M.
    Basiri J.
    Seyedhossein L.
    Shakery A.
    2010 5th International Symposium on Telecommunications, IST 2010, 2010, : 531 - 536
  • [3] A case-based technique for tracking concept drift in spam filtering
    Delany, SJ
    Cunningham, P
    Tsymbal, A
    Coyle, L
    APPLICATIONS AND INNOVATIONS IN INTELLIGENT SYSTEMS XII, PROCEEDINGS, 2005, : 3 - 16
  • [4] Tracking concept drift at feature selection stage in SpamHunting:: An anti-spam instance-based reasoning system
    Mendez, J. R.
    Fdez-Riverola, F.
    Iglesias, E. L.
    Diaz, F.
    Corchado, J. M.
    ADVANCES IN CASE-BASED REASONING, PROCEEDINGS, 2006, 4106 : 504 - 518
  • [5] A case-based technique for tracking concept drift in spam filtering
    Delany, SJ
    Cunningham, P
    Tsymbal, A
    Coyle, L
    KNOWLEDGE-BASED SYSTEMS, 2005, 18 (4-5) : 187 - 195
  • [6] Feature-based analyses of concept drift
    Hinder, Fabian
    Vaquet, Valerie
    Hammer, Barbara
    NEUROCOMPUTING, 2024, 600
  • [7] Genetic-based Feature Selection for Spam Detection
    Arani, Seyyed Hossein Seyyedi
    Mozaffari, Saeed
    2013 21ST IRANIAN CONFERENCE ON ELECTRICAL ENGINEERING (ICEE), 2013,
  • [8] Concentration Based Feature Construction Approach for Spam Detection
    Tan, Ying
    Deng, Chao
    Ruan, Guangchen
    IJCNN: 2009 INTERNATIONAL JOINT CONFERENCE ON NEURAL NETWORKS, VOLS 1- 6, 2009, : 510 - 515
  • [9] The Entropy-Based Time Domain Feature Extraction for Online Concept Drift Detection
    Ding, Fengqian
    Luo, Chao
    ENTROPY, 2019, 21 (12)
  • [10] A Novel Framework for Spam Hunting by Tracking Concept Drift
    Bindu, V
    Thomas, Ciza
    SECOND INTERNATIONAL CONFERENCE ON COMPUTER NETWORKS AND COMMUNICATION TECHNOLOGIES, ICCNCT 2019, 2020, 44 : 918 - 926