Mining California vital statistical data

被引:1
|
作者
Zhang, Du [1 ]
Quoc Luan Ha [1 ]
Lu, Meiliu [1 ]
机构
[1] Calif State Univ, Dept Comp Sci, Sacramento, CA 95819 USA
关键词
vital statistics data; causes of death; data mining; predictive models; Cubist;
D O I
10.1504/IJCAT.2006.011999
中图分类号
TP39 [计算机的应用];
学科分类号
081203 ; 0835 ;
摘要
Vital statistics data offer a fertile ground for data mining. In this paper, we discuss the results of a data-mining project on the causes of death aspect of the vital statistics data in the state of California. A data-mining tool called Cubist is used to build predictive models out of two million cases over a nine-year period. The objective of our study is to discover knowledge (trends, correlations or patterns) that may not be gleaned through standard techniques. The generated predictive models allow pertinent state agencies to gain insight into various aspects of the death rates in the state of California, to predict health issues related to the causes of death, to offer an aid to decision - or policy-making process and to provide useful information services to the customers. The results obtained in our study contain valuable new information.
引用
收藏
页码:281 / 297
页数:17
相关论文
共 50 条
  • [1] Mining California vital statistics data
    Zhang, D
    Ha, QL
    Lu, ML
    2001 IEEE INTERNATIONAL CONFERENCE ON DATA MINING, PROCEEDINGS, 2001, : 671 - 672
  • [2] Statistical data mining
    Banks, David L.
    WILEY INTERDISCIPLINARY REVIEWS-COMPUTATIONAL STATISTICS, 2010, 2 (01): : 9 - 25
  • [3] Statistical models for data mining
    Giudici, P
    Heckerman, D
    Whittaker, J
    DATA MINING AND KNOWLEDGE DISCOVERY, 2001, 5 (03) : 163 - 165
  • [4] Statistical inference and data mining
    Glymour, C
    Madigan, D
    Pregibon, D
    Smyth, P
    COMMUNICATIONS OF THE ACM, 1996, 39 (11) : 35 - 41
  • [5] Statistical Models for Data Mining
    Paolo Giudici
    David Heckerman
    Joe Whittaker
    Data Mining and Knowledge Discovery, 2001, 5 : 163 - 165
  • [6] A statistical perspective on data mining
    Hosking, JRM
    Pednault, EPD
    Sudan, M
    FUTURE GENERATION COMPUTER SYSTEMS-THE INTERNATIONAL JOURNAL OF ESCIENCE, 1997, 13 (2-3): : 117 - 134
  • [7] Statistical perspective on data mining
    IBM T.J. Watson Research Cent, Yorktown Heights, United States
    Future Gener Comput Syst, 2-3 (117-134):
  • [8] Data mining and statistical analysis of educational data
    Brites, Nuno M.
    Melgueira, Pedro
    Rodrigues, Irene P.
    Ferreira, Ligia
    INTERNATIONAL JOURNAL OF APPLIED MATHEMATICS & STATISTICS, 2018, 57 (05): : 103 - 114
  • [9] Data mining, statistical methods mining, and history of statistics
    Parzen, E
    MINING AND MODELING MASSIVE DATA SETS IN SCIENCE, ENGINEERING, AND BUSINESS WITH A SUBTHEME IN ENVIRONMENTAL STATISTICS, 1997, 29 (01): : 365 - 374
  • [10] Statistical Themes and Lessons for Data Mining
    Clark Glymour
    David Madigan
    Daryl Pregibon
    Padhraic Smyth
    Data Mining and Knowledge Discovery, 1997, 1 : 11 - 28