Automatic discovery of the sequential accesses from web log data files via a genetic algorithm

被引:11
|
作者
Tug, Emine [1 ]
Sakiroglu, Merve [1 ]
Arslan, Ahmet [1 ]
机构
[1] Selcuk Univ, Dept Comp Sci, Konya 42300, Turkey
关键词
web mining; genetic algorithm; knowledge discovery; sequential access;
D O I
10.1016/j.knosys.2005.10.008
中图分类号
TP18 [人工智能理论];
学科分类号
081104 ; 0812 ; 0835 ; 1405 ;
摘要
This paper is concerned with finding sequential accesses from web log files, using 'Genetic Algorithm' (GA). Web log files are independent from servers, and they are ASCII format. Each transaction, whether completed or not, is recorded in the web log files and these files are unstructured for knowledge discovery in database techniques. Data which is stored in web logs have become important for discovering of user behaviors since the using of internet increased rapidly. Analyzing of these log files is one of the important research area of web mining. Especially, with the advent of CRM (Customer Resource Management) issues in business circle, most of the modem firms operating web sites for several purposes are now adopting web-mining as a strategic way of capturing knowledge about potential needs of target customers, future trends in the market and other management factors. Our work (ALMG-Automatic Log Mining via Genetic) has mined web log files via genetic algorithm. When we search the studies about web mining in literature, it can be seen that, GA is generally used in web content and web structure mining. On the other hand, ALMG is a study about web mining usage. The difference between ALMG and other similar works at literature is this point. As for in another work that we are encountering, GA is used for processing the data between HTML tags which are placed at client PC. But ALMG extracts information from data which is placed at server. It is thought to use log files is an advantage for our purpose. Because, we find the character of requests which is made to the server than detect a single person's behavior. We developed an application with this purpose. Firstly, the application is analyzed web log files, than found sequential accessed page groups automatically. (c) 2005 Elsevier B.V. All rights reserved.
引用
收藏
页码:180 / 186
页数:7
相关论文
共 50 条
  • [21] Study on data preprocessing algorithm in web log mining
    Yuan, F
    Wang, LJ
    Yu, G
    2003 INTERNATIONAL CONFERENCE ON MACHINE LEARNING AND CYBERNETICS, VOLS 1-5, PROCEEDINGS, 2003, : 28 - 32
  • [22] Automatic calculation of students' conceptions in elementary algebra from Aplusix log files
    Nicaud, Jean-Francois
    Chaachoua, Hamid
    Bittar, Marilena
    INTELLIGENT TUTORING SYSTEMS, PROCEEDINGS, 2006, 4053 : 433 - 442
  • [23] Model-Driven Development of Multidimensional Models from Web Log Files
    Hernandez, Paul
    Garrigos, Irene
    Mazon, Jose-Norberto
    ADVANCES IN CONCEPTUAL MODELING: APPLICATIONS AND CHALLENGES, 2010, 6413 : 170 - 179
  • [24] An Interactive Web-Based Toolset for Knowledge Discovery from Short Text Log Data
    Stewart, Michael
    Liu, Wei
    Cardell-Oliver, Rachell
    Griffin, Mark
    ADVANCED DATA MINING AND APPLICATIONS, ADMA 2017, 2017, 10604 : 853 - 858
  • [25] Automatic discovery of synonyms and lexicalizations from the Web
    Sanchez, David
    Moreno, Antonio
    ARTIFICIAL INTELLIGENCE RESEARCH AND DEVELOPMENT, 2005, 131 : 205 - 212
  • [26] Automatic information discovery from the "Invisible web"
    Lin, KI
    Chen, H
    INTERNATIONAL CONFERENCE ON INFORMATION TECHNOLOGY: CODING AND COMPUTING, PROCEEDINGS, 2002, : 332 - 337
  • [27] A Kind of Improved Data Clustering Algorithm in Web Log Mining
    Guo, Jin
    Zhang, Shengbing
    Qiu, Zheng
    PROCEEDINGS OF THE 2015 INTERNATIONAL CONFERENCE ON INTELLIGENT SYSTEMS RESEARCH AND MECHATRONICS ENGINEERING, 2015, 121 : 2115 - 2119
  • [28] Recovering Deleted Browsing Artifacts from Web Browser Log Files in Linux Environment
    Anuradha, P.
    Kumar, Raj T.
    Sobhana, N. V.
    2016 SYMPOSIUM ON COLOSSAL DATA ANALYSIS AND NETWORKING (CDAN), 2016,
  • [29] DISCOVERY OF WEB PATTERN FROM WEB LOGS FILES USING ENHANCED GRAPH GRAMMAR APPROACH
    Anbarasi, M. S.
    Vasanthi, B.
    2017 INTERNATIONAL CONFERENCE ON INFORMATION COMMUNICATION AND EMBEDDED SYSTEMS (ICICES), 2017,
  • [30] Reliable Biomarker discovery from Metagenomic data via RegLRSD algorithm
    Alshawaqfeh, Mustafa
    Bashaireh, Ahmad
    Serpedin, Erchin
    Suchodolski, Jan
    BMC BIOINFORMATICS, 2017, 18