• DocumentCode
    3125965
  • Title

    Lessons learned at 208K: Towards debugging millions of cores

  • Author

    Lee, Gregory L. ; Ahn, Dong H. ; Arnold, Dorian C. ; De Supinski, Bronis R. ; Legendre, Matthew ; Miller, Barton P. ; Schulz, Martin ; Liblit, Ben

  • Author_Institution
    Comput. Directorate, Lawrence Livermore Nat. Lab., Livermore, CA, USA
  • fYear
    2008
  • fDate
    15-21 Nov. 2008
  • Firstpage
    1
  • Lastpage
    9
  • Abstract
    Petascale systems will present several new challenges to performance and correctness tools. Such machines may contain millions of cores, requiring that tools use scalable data structures and analysis algorithms to collect and to process application data. In addition, at such scales, each tool itself will become a large parallel application - already, debugging the full Blue-Gene/L (BG/L) installation at the Lawrence Livermore National Laboratory requires employing 1664 tool daemons. To reach such sizes and beyond, tools must use a scalable communication infrastructure and manage their own tool processes efficiently. Some system resources, such as the file system, may also become tool bottlenecks. In this paper, we present challenges to petascale tool development, using the stack trace analysis tool (STAT) as a case study. STAT is a lightweight tool that gathers and merges stack traces from a parallel application to identify process equivalence classes. We use results gathered at thousands of tasks on an Infiniband cluster and results up to 208 K processes on BG/L to identify current scalability issues as well as challenges that will be faced at the petascale. We then present implemented solutions to these challenges and show the resulting performance improvements. We also discuss future plans to meet the debugging demands of petascale machines.
  • Keywords
    parallel machines; parallel programming; program debugging; program diagnostics; software tools; Blue-Gene/L installation; STAT; correctness tool; debugging; performance tool; petascale systems; petascale tool development; scalable communication infrastructure; scalable data analysis; scalable data structures; stack trace analysis tool; tool daemons; Algorithm design and analysis; Application software; Data analysis; Data structures; Debugging; File systems; Laboratories; Large-scale systems; Scalability; System software;
  • fLanguage
    English
  • Publisher
    ieee
  • Conference_Titel
    High Performance Computing, Networking, Storage and Analysis, 2008. SC 2008. International Conference for
  • Conference_Location
    Austin, TX
  • Print_ISBN
    978-1-4244-2834-2
  • Electronic_ISBN
    978-1-4244-2835-9
  • Type

    conf

  • DOI
    10.1109/SC.2008.5218557
  • Filename
    5218557