Autonomic failure prediction based on manifold learning for large-scale distributed systems

被引:6
|
作者
Lu X. [1 ]
Wang H.-Q. [1 ]
Zhou R.-J. [1 ]
Ge B.-Y. [1 ]
机构
[1] College of Computer Science and Technology, Harbin Engineering University
基金
中国国家自然科学基金;
关键词
autonomic computing; failure prediction; locally linear embedding; manifold learning;
D O I
10.1016/S1005-8885(09)60497-0
中图分类号
学科分类号
摘要
This article investigates autonomic failure prediction in large-scale distributed systems with nonlinear dimensionality reduction to automatically extract failure features. Most existing methods for failure prediction focus on building prediction models or heuristic rules by discovering failure patterns, but the process of feature extraction before failure patterns recognition is rarely considered due to the increasing complexity of modern distributed systems. In this work, a novel performance-centric approach to automate failure prediction is proposed based on manifold learning (ML). In addition, the ML algorithm named supervised locally linear embedding (SLLE) is applied to achieve feature extraction. To generalize the dimensionality reduction mapping, the nonlinear mapping approximation and optimization solution is also proposed. In experimental work a file transfer test bed with fault injection is developed which can gather multilevel performance metrics transparently. Based on the runtime monitoring of these metrics, the SLLE method can automatically predict more than 50 of the central processing unit (CPU) and memory failures, and around 70 of the network failure. © 2010 The Journal of China Universities of Posts and Telecommunications.
引用
收藏
页码:116 / 124
页数:8
相关论文
共 50 条
  • [21] Stability of large-scale distributed parameter systems
    Ladde, GS
    Li, TT
    DYNAMIC SYSTEMS AND APPLICATIONS, 2002, 11 (03): : 311 - 323
  • [22] Energy efficiency in large-scale distributed systems
    Tuan Anh Trinh
    Hlavacs, Helmut
    Talia, Domenico
    FUTURE GENERATION COMPUTER SYSTEMS-THE INTERNATIONAL JOURNAL OF GRID COMPUTING AND ESCIENCE, 2012, 28 (05): : 743 - 744
  • [23] Monitoring and control of large-scale distributed systems
    Legrand, C.
    GRID AND CLOUD COMPUTING: CONCEPTS AND PRACTICAL APPLICATIONS, 2016, 192 : 101 - 151
  • [24] Independent recovery in large-scale distributed systems
    Triantafillou, P
    IEEE TRANSACTIONS ON SOFTWARE ENGINEERING, 1996, 22 (11) : 812 - 826
  • [25] Expectations and challenges in large-scale distributed systems
    Bacon, J
    IEEE CONCURRENCY, 2000, 8 (01): : 2 - 3
  • [26] A dependability layer for large-scale distributed systems
    Cristea, Valentin
    Dobre, C.
    Pop, F.
    Stratan, C.
    Costan, A.
    Leordeanu, C.
    Tirsa, E.
    INTERNATIONAL JOURNAL OF GRID AND UTILITY COMPUTING, 2011, 2 (02) : 109 - 118
  • [27] Large-scale network intrusion detection based on distributed learning algorithm
    Daxin Tian
    Yanheng Liu
    Yang Xiang
    International Journal of Information Security, 2009, 8 : 25 - 35
  • [28] Distributed Orchestration in Large-scale IoT Systems
    Yigitoglu, Emre
    Liu, Ling
    Looper, Margaret
    Pu, Calton
    2017 IEEE 2ND INTERNATIONAL CONGRESS ON INTERNET OF THINGS (IEEE ICIOT), 2017, : 58 - 65
  • [29] Robust Scheduling for Large-Scale Distributed Systems
    Lee, Young Choon
    King, Jayden
    Kim, Young Ki
    Hong, Seok-Hee
    2020 IEEE 19TH INTERNATIONAL CONFERENCE ON TRUST, SECURITY AND PRIVACY IN COMPUTING AND COMMUNICATIONS (TRUSTCOM 2020), 2020, : 38 - 45
  • [30] Adaptation Engine for Large-Scale Distributed Systems
    Nemes, Tania
    COMPUTER AIDED SYSTEMS THEORY - EUROCAST 2015, 2015, 9520 : 244 - 251