• DocumentCode
    580398
  • Title

    Failure analysis of distributed scientific workflows executing in the cloud

  • Author

    Samak, Taghrid ; Gunter, Dan ; Goode, Monte ; Deelman, Ewa ; Juve, Gideon ; Silva, Fabio ; Vahi, Karan

  • Author_Institution
    Lawrence Berkeley Nat. Lab., Berkeley, CA, USA
  • fYear
    2012
  • fDate
    22-26 Oct. 2012
  • Firstpage
    46
  • Lastpage
    54
  • Abstract
    This work presents models characterizing failures observed during the execution of large scientific applications on Amazon EC2. Scientific workflows are used as the underlying abstraction for application representations. As scientific workflows scale to hundreds of thousands of distinct tasks, failures due to software and hardware faults become increasingly common. We study job failure models for data collected from 4 scientific applications, by our Stampede framework. In particular, we show that a Naive Bayes classifier can accurately predict the failure probability of jobs. The models allow us to predict job failures for a given execution resource and then use these failure predictions for two higher-level goals: (1) to suggest a better job assignment, and (2) to provide quantitative feedback to the workflow component developer about the robustness of their application codes.
  • Keywords
    Bayes methods; cloud computing; failure analysis; natural sciences computing; Amazon EC2; Stampede framework; cloud computing; distributed scientific workflows; failure analysis; hardware faults; job assignment; job failure; naive Bayes classifier; software faults; Accuracy; Broadband communication; Data models; Engines; Erbium; Failure analysis; Predictive models;
  • fLanguage
    English
  • Publisher
    ieee
  • Conference_Titel
    Network and service management (cnsm), 2012 8th international conference and 2012 workshop on systems virtualiztion management (svm)
  • Conference_Location
    Las Vegas, NV
  • Print_ISBN
    978-1-4673-3134-0
  • Electronic_ISBN
    978-3-901882-48-7
  • Type

    conf

  • Filename
    6379991