• DocumentCode
    1933832
  • Title

    Locality aware concurrent start for stencil applications

  • Author

    Shrestha, Sunil ; Gao, Guang R. ; Manzano, Joseph ; Marquez, Andres ; Feo, John

  • fYear
    2015
  • fDate
    7-11 Feb. 2015
  • Firstpage
    157
  • Lastpage
    166
  • Abstract
    Stencil computations are at the heart of many physical simulations used in scientific codes. Thus, there exists a plethora of optimization efforts for this family of computations. Among these techniques, tiling techniques that allow concurrent start have proven to be very efficient in providing better performance for these critical kernels. Nevertheless, with many core designs being the norm, these optimization techniques might not be able to fully exploit locality (both spatial and temporal) on multiple levels of the memory hierarchy without compromising parallelism. It is no longer true that the machine can be seen as a homogeneous collection of nodes with caches, main memory and an interconnect network. New architectural designs exhibit complex grouping of nodes, cores, threads, caches and memory connected by an ever evolving network-on-chip design. These new designs may benefit greatly from carefully crafted schedules and groupings that encourage parallel actors (i.e. threads, cores or nodes) to be aware of the computational history of other actors in close proximity. In this paper, we provide an efficient tiling technique that allows hierarchical concurrent start for memory hierarchy aware tile groups. Each execution schedule and tile shape exploit the available parallelism, load balance and locality present in the given applications. We demonstrate our technique on the Intel Xeon Phi architecture with selected and representative stencil kernels. We show improvement ranging from 5.58% to 31.17% over existing state-of-the-art techniques.
  • Keywords
    mobile computing; network-on-chip; optimisation; storage management chips; Intel Xeon Phi architecture; interconnect network; locality aware concurrent start; memory hierarchy aware tile groups; network-on-chip design; optimization techniques; stencil computations; tiling techniques; Diamonds; Heating; Instruction sets; Kernel; Optimization; Parallel processing; Pluto;
  • fLanguage
    English
  • Publisher
    ieee
  • Conference_Titel
    Code Generation and Optimization (CGO), 2015 IEEE/ACM International Symposium on
  • Conference_Location
    San Francisco, CA
  • Type

    conf

  • DOI
    10.1109/CGO.2015.7054196
  • Filename
    7054196