• DocumentCode
    167439
  • Title

    Scalable Critical Path Analysis for Hybrid MPI-CUDA Applications

  • Author

    Schmitt, Felix ; Dietrich, Robert ; Juckeland, Guido

  • Author_Institution
    Center for Inf. Services & High Performance Comput. (ZIH), Tech. Univ. Dresden, Dresden, Germany
  • fYear
    2014
  • fDate
    19-23 May 2014
  • Firstpage
    908
  • Lastpage
    915
  • Abstract
    Utilizing accelerators in heterogeneous systems is an established approach for designing peta-scale applications. Today, CUDA offers a rich programming interface for GPU accelerators but requires developers to incorporate several layers of parallelism on both CPU and GPU. From this increasing program complexity emerges the need for sophisticated performance tools. This work contributes by analyzing hybrid MPI-CUDA programs for their critical path, a property proven to effectively identify application bottlenecks. We developed a tool which constructs a dependency graph based on an execution trace and the inherent dependencies of the programming models CUDA and MPI. Thereafter, it detects wait-states and attributes blame to responsible activities. Together with the property of being on the critical path we can identify activities that are most viable for optimization. The developed approach has been demonstrated with suitable examples to be both scalable and correct. Furthermore, we establish a new categorization of CUDA inefficiency patterns ensuing from the dependencies between CUDA activities.
  • Keywords
    critical path analysis; message passing; optimisation; parallel processing; HPC; critical path analysis; dependency graph; execution trace; high-performance computing; hybrid MPI-CUDA applications; performance optimization; programming models; Graphics processing units; Kernel; Optimization; Performance analysis; Runtime; Synchronization; CUDA; GPGPU; MPI; critical path analysis; performance analysis; performance optimization; wait-states;
  • fLanguage
    English
  • Publisher
    ieee
  • Conference_Titel
    Parallel & Distributed Processing Symposium Workshops (IPDPSW), 2014 IEEE International
  • Conference_Location
    Phoenix, AZ
  • Print_ISBN
    978-1-4799-4117-9
  • Type

    conf

  • DOI
    10.1109/IPDPSW.2014.103
  • Filename
    6969479