DocumentCode :
3227670
Title :
A Survey on Failure Prediction of Large-Scale Server Clusters
Author :
Xue, Zhenghua ; Dong, Xiaoshe ; Ma, Siyuan ; Dong, Weiqing
Author_Institution :
Xi´´an Jiaotong Univ., Xian
Volume :
2
fYear :
2007
fDate :
July 30 2007-Aug. 1 2007
Firstpage :
733
Lastpage :
738
Abstract :
As the size and complexity of cluster systems grows, failure rates accelerate dramatically. To reduce the disaster caused by failures, it is desirable to identify the potential failures ahead of their occurrence. In this paper, we survey the state of the art in failure prediction of cluster systems. The characteristic of failures in cluster systems are addressed, and some statistic results are shown. We explore the ways of the collection and preprocessing of data for failure prediction, and suggest a procedure for preprocessing the records in automatically generated log files. Focused on the main idea of five prediction methods, including statistic based threshold, time series analysis, rule-based classification, Bayesian network models and semi-Markov process models, are analyzed respectively. In addition, concerning the accuracy and practicality, we present five metrics for evaluating the failure prediction techniques and compare the five techniques with the five metrics.
Keywords :
Bayes methods; Markov processes; statistical analysis; time series; workstation clusters; Bayesian network model; failure prediction; large-scale server cluster system; rule-based classification; semiMarkov process model; statistic based threshold; time series analysis; Artificial intelligence; Hardware; High performance computing; Large-scale systems; Predictive models; Redundancy; Software engineering; Statistical distributions; Statistics; Time series analysis;
fLanguage :
English
Publisher :
ieee
Conference_Titel :
Software Engineering, Artificial Intelligence, Networking, and Parallel/Distributed Computing, 2007. SNPD 2007. Eighth ACIS International Conference on
Conference_Location :
Qingdao
Print_ISBN :
978-0-7695-2909-7
Type :
conf
DOI :
10.1109/SNPD.2007.284
Filename :
4287779
Link To Document :
بازگشت