• DocumentCode
    2990076
  • Title

    Software prefetch on core micro-architecture applied to irregular codes

  • Author

    Ammenouche, Samir ; Singh, David E. ; Carretero, Jesús ; Jalby, William

  • Author_Institution
    Univ. of Versailles, Versailles, France
  • fYear
    2011
  • fDate
    4-8 July 2011
  • Firstpage
    264
  • Lastpage
    272
  • Abstract
    Sparse computations constitute one of the most important areas of numerical algebra and scientific computing. Be cause of indirect addressing, sparse codes exhibit irregular patterns of references to memory. While there are many studies performing high level optimizations on sparse computing, few deal with software prefetch. This is due to the irregular memory accesses which are incompatible with regular prefetch and it is due to the high efficiency and complexity of hardware prefetch units included in modern processors, like the Intel Core micro-architecture. In this paper, we show the efficiency and the limitations of hardware prefetch units, and we propose a technique to use software prefetch instructions in combination with hardware support to better manage cache and improve the overall code performance. To achieve this goal, the cache behavior of the sparse matrix vector multiplication (SpMV) is analyzed focusing on the code structure and the sequence order of the data. Main cache parameters are identified and their impact on the cache performance is evaluated. These parameters are included in a matrix analyzer to determine in advance the efficiency of the software prefetch. Furthermore, the software prefetch efficiency is analyzed on a large set of sparse matrices. Experimental results show an accurate prediction of the matrix analyzer and a maximum improvement of 40% in the execution time.
  • Keywords
    matrix multiplication; microprocessor chips; software engineering; storage management; Intel Core microarchitecture; cache behavior; code performance; hardware prefetch unit; matrix analyzer; software prefetch; sparse code; sparse computing; sparse matrix; sparse matrix vector multiplication; Arrays; Hardware; Kernel; Optimization; Prefetching; Sparse matrices; Intel Core; L2 cache; software prefetch; sparse matrix;
  • fLanguage
    English
  • Publisher
    ieee
  • Conference_Titel
    High Performance Computing and Simulation (HPCS), 2011 International Conference on
  • Conference_Location
    Istanbul
  • Print_ISBN
    978-1-61284-380-3
  • Type

    conf

  • DOI
    10.1109/HPCSim.2011.5999833
  • Filename
    5999833