• DocumentCode
    1799916
  • Title

    COMP: Compiler Optimizations for Manycore Processors

  • Author

    Linhai Song ; Min Feng ; Ravi, Nishkam ; Yi Yang ; Chakradhar, Srimat

  • Author_Institution
    Comput. Sci. Dept., Univ. of Wisconsin-Madison, Madison, WI, USA
  • fYear
    2014
  • fDate
    13-17 Dec. 2014
  • Firstpage
    659
  • Lastpage
    671
  • Abstract
    Applications executing on multicore processors can now easily offload computations to many core processors, such as Intel Xeon Phi coprocessors. However, it requires high levels of expertise and effort to tune such offloaded applications to realize high-performance execution. Previous efforts have focused on optimizing the execution of offloaded computations on many core processors. However, we observe that the data transfer overhead between multicore and many core processors, and the limited device memories of many core processors often constrain the performance gains that are possible by offloading computations. In this paper, we present three source-to-source compiler optimizations that can significantly improve the performance of applications that offload computations to many core processors. The first optimization automatically transforms offloaded codes to enable data streaming, which overlaps data transfer between multicore and many core processors with computations on these processors to hide data transfer overhead. This optimization is also designed to minimize the memory usage on many core processors, while achieving the optimal performance. The second compiler optimization re-orders computations to regularize irregular memory accesses. It enables data streaming and factorization on many core processors, even when the memory access patterns in the original source codes are irregular. Finally, our new shared memory mechanism provides efficient support for transferring large pointer-based data structures between hosts and many core processors. Our evaluation shows that the proposed compiler optimizations benefit 9 out of 12 benchmarks. Compared with simply offloading the original parallel implementations of these benchmarks, we can achieve 1.16x-52.21x speedups.
  • Keywords
    data structures; optimising compilers; shared memory systems; source code (software); storage management; compiler optimization for manycore processors; data streaming; data transfer overhead hiding; memory access patterns; memory usage minimization; multicore processors; offloaded codes; pointer-based data structures; shared memory mechanism; source-to-source compiler optimizations; vectorization; Benchmark testing; Coprocessors; Data transfer; Microwave integrated circuits; Multicore processing; Optimization; Program processors; Intel MIC; compiler optimizations; manycore coprocessors; offload;
  • fLanguage
    English
  • Publisher
    ieee
  • Conference_Titel
    Microarchitecture (MICRO), 2014 47th Annual IEEE/ACM International Symposium on
  • Conference_Location
    Cambridge
  • ISSN
    1072-4451
  • Type

    conf

  • DOI
    10.1109/MICRO.2014.30
  • Filename
    7011425