• DocumentCode
    2866865
  • Title

    A Runtime Fault Detection Method for HPC Cluster

  • Author

    Linping, Wu ; Hongbing, Luo ; Jianfeng, Zhan ; Dan, Meng

  • Author_Institution
    HPCC, Inst. of Appl. Phys. & Comput. Mathematic, Beijing, China
  • fYear
    2011
  • fDate
    20-22 Oct. 2011
  • Firstpage
    68
  • Lastpage
    72
  • Abstract
    As the number of nodes keeps increasing, faults have become commonplace for HPC cluster. For fast recovery from faults, the fault detection method is necessary. Based on the usage patterns of HPC cluster, a automatic runtime fault detection mechanism is proposed in this paper: First, the normal activities for nodes in HPC cluster are modeled using runtime state by clustering analysis, Second, the fault detection process is implemented by comparing the current runtime state of nodes with normal activity models. A fault alarm is made immediately when the current runtime state deviates from the normal activity models. In the experiments, the faults are simulated by fault injection methods and the experimental results show that the runtime fault detection method in this paper can detect faults with high accuracy.
  • Keywords
    fault diagnosis; parallel processing; pattern clustering; HPC cluster analysis; automatic runtime fault detection mechanism; fault alarm; fault injection method; fault recovery; normal activity model; runtime state; Aging; Computational modeling; Fault detection; Runtime; Software; System recovery; Vectors; HPC cluster; Runtime fault detection; SOM;
  • fLanguage
    English
  • Publisher
    ieee
  • Conference_Titel
    Parallel and Distributed Computing, Applications and Technologies (PDCAT), 2011 12th International Conference on
  • Conference_Location
    Gwangju
  • Print_ISBN
    978-1-4577-1807-6
  • Type

    conf

  • DOI
    10.1109/PDCAT.2011.9
  • Filename
    6118961