• DocumentCode
    1858119
  • Title

    Optimizing Process-to-Core Mappings for Two Dimensional Broadcast/Reduce on Multicore Architectures

  • Author

    Karlsson, Christer ; Davies, Teresa ; Ding, Chong ; Liu, Hui ; Chen, Zizhong

  • Author_Institution
    Dept. of Math. & Comput. Sci., Colorado Sch. of Mines, Golden, CO, USA
  • fYear
    2011
  • fDate
    13-16 Sept. 2011
  • Firstpage
    404
  • Lastpage
    413
  • Abstract
    In today´s high performance computing, many MPI programs (e.g., ScaLAPACK applications, High Performance Linpack Benchmark HPL, and many PDE solvers based on domain decomposition methods) organize their computational processes as multidimensional process grids. Communications are often necessary in each dimension. Multidimensional broadcast, where a broadcast has to be performed in each dimension, is one of the many operations in applications that use multidimensional process grids. In this paper, we study the impact of the MPI process-to-core mapping on the performance of multidimensional broadcast operations. We show that the default process-to-core mappings in today´s state-of-the-art MPI implementations are often sub-optimal for multidimensional broadcast. We propose an application-level multicore-aware process-to-core re-mapping scheme that is capable of achieving optimal performance for multidimensional broadcast operations. The proposed multicore-aware process-to-core re-mapping scheme improves the performance of multidimensional broadcast operations by up to 64% over the default mapping scheme on the world´s current eighth fastest supercomputer, Kraken, at the Oak Ridge National Laboratory.
  • Keywords
    application program interfaces; message passing; multiprocessing systems; parallel machines; Kraken; MPI process to core mapping; MPI program; Oak Ridge National Laboratory; application level multicore-aware process to core remapping scheme; message passing interface; multidimensional broadcast performance; multidimensional process grid; optimal performance; supercomputer; Hardware; Multicore processing; Optimization; Pipelines; Tiles; Topology; Cartesian Topology; Collective Communication; Message Passing Interface (MPI); Multicore; Process-to-Core Mapping;
  • fLanguage
    English
  • Publisher
    ieee
  • Conference_Titel
    Parallel Processing (ICPP), 2011 International Conference on
  • Conference_Location
    Taipei City
  • ISSN
    0190-3918
  • Print_ISBN
    978-1-4577-1336-1
  • Electronic_ISBN
    0190-3918
  • Type

    conf

  • DOI
    10.1109/ICPP.2011.26
  • Filename
    6047208