Data mining using relational database management systems

被引:0
|
作者
Zou, B [1 ]
Ma, X
Kemme, B
Newton, G
Precup, D
机构
[1] McGill Univ, Montreal, PQ, Canada
[2] Natl Res Council Canada, Ottawa, ON K1A 0R6, Canada
关键词
D O I
暂无
中图分类号
TP18 [人工智能理论];
学科分类号
081104 ; 0812 ; 0835 ; 1405 ;
摘要
Software packages providing a whole set of data mining and machine learning algorithms are attractive because they allow experimentation with many kinds of algorithms in an easy setup. However, these packages are often based on main-memory data structures, limiting the amount of data they can handle. In this paper we use a relational database as secondary storage in order to eliminate this limitation. Unlike existing approaches, which often focus on optimizing a single algorithm to work with a database backend, we propose a general approach, which provides a database interface for several algorithms at once. We have taken a popular machine learning software package, Weka, and added a relational storage manager as back-tier to the system. The extension is transparent to the algorithms implemented in Weka, since it is hidden behind Weka's standard main-memory data structure interface. Furthermore, some general mining tasks are transfered into the database system to speed up execution. We tested the extended system, refered to as WekaDB, and our results show that it achieves a much higher scalability than Weka, while providing the same output and maintaining good computation time.
引用
收藏
页码:657 / 667
页数:11
相关论文
共 50 条
  • [1] Handling Big Data in Relational Database Management Systems
    ElDahshan, Kamal
    Selim, Eman
    Ebada, Ahmed Ismail
    Abouhawwash, Mohamed
    Nam, Yunyoung
    Behery, Gamal
    [J]. CMC-COMPUTERS MATERIALS & CONTINUA, 2022, 72 (03): : 5149 - 5164
  • [2] Data mining support in database management systems
    Morzy, T
    Wojciechowski, M
    Zakrzewicz, M
    [J]. DATA WAREHOUSING AND KNOWLEDGE DISCOVERY, PROCEEDINGS, 2000, 1874 : 382 - 392
  • [3] Extension of relational management systems with data mining capabilities
    Martinez, JF
    Wasilewska, A
    Hadjimichael, M
    Fernandez, C
    Menasalvas, E
    [J]. ROUGH SETS AND CURRENT TRENDS IN COMPUTING, PROCEEDINGS, 2002, 2475 : 421 - 424
  • [4] Managing method of spatial data in relational database management systems
    Zhu, Tie-Wen
    Zhong, Zhi-Nong
    Jing, Ning
    [J]. Ruan Jian Xue Bao/Journal of Software, 2002, 13 (01): : 8 - 14
  • [5] Implementing Multi-relational Mining with Relational Database Systems
    Inuzuka, Nobuhiro
    Makino, Toshiyuki
    [J]. KNOWLEDGE-BASED AND INTELLIGENT INFORMATION AND ENGINEERING SYSTEMS, PT II, PROCEEDINGS, 2009, 5712 : 672 - 680
  • [6] A layered optimizer for mining association rules over relational database management systems
    Dudgikar, M
    Chakravarthy, S
    Liuzzi, R
    Wong, L
    [J]. IKE'03: PROCEEDINGS OF THE INTERNATIONAL CONFERENCE ON INFORMATION AND KNOWLEDGE ENGINEERING, VOLS 1 AND 2, 2003, : 422 - 427
  • [7] Cryptography and relational database management systems
    He, JM
    Wang, M
    [J]. 2001 INTERNATIONAL DATABASE ENGINEERING & APPLICATIONS SYMPOSIUM, PROCEEDINGS, 2001, : 273 - 284
  • [8] FORMAL RESOURCE DATA MODEL IN RELATIONAL DATABASE-MANAGEMENT SYSTEMS
    KRAMARENKO, RP
    GOLOSHCHUK, IA
    [J]. CYBERNETICS, 1984, 20 (05): : 697 - 703
  • [9] Efficient integration of data mining techniques in database management systems
    Bentayeb, F
    Darmont, J
    Udréa, C
    [J]. INTERNATIONAL DATABASE ENGINEERING AND APPLICATIONS SYMPOSIUM, PROCEEDINGS, 2004, : 59 - 67
  • [10] Data mining of inclusion dependency from relational database
    Wang, SL
    Chen, YC
    Hong, TP
    [J]. KNOWLEDGE-BASED INTELLIGENT INFORMATION ENGINEERING SYSTEMS & ALLIED TECHNOLOGIES, PTS 1 AND 2, 2001, 69 : 505 - 509