• DocumentCode
    2241233
  • Title

    Many-Thread Aware Prefetching Mechanisms for GPGPU Applications

  • Author

    Lee, Jaekyu ; Lakshminarayana, Nagesh B. ; Kim, Hyesoon ; Vuduc, Richard

  • Author_Institution
    Coll. of Comput., Georgia Inst. of Technol., Atlanta, GA, USA
  • fYear
    2010
  • fDate
    4-8 Dec. 2010
  • Firstpage
    213
  • Lastpage
    224
  • Abstract
    We consider the problem of how to improve memory latency tolerance in massively multithreaded GPGPUs when the thread-level parallelism of an application is not sufficient to hide memory latency. One solution used in conventional CPU systems is prefetching, both in hardware and software. However, we show that straightforwardly applying such mechanisms to GPGPU systems does not deliver the expected performance benefits and can in fact hurt performance when not used judiciously. This paper proposes new hardware and software prefetching mechanisms tailored to GPGPU systems, which we refer to as many-thread aware prefetching (MT-prefetching) mechanisms. Our software MT-prefetching mechanism, called inter-thread prefetching, exploits the existence of common memory access behavior among fine-grained threads. For hardware MT-prefetching, we describe a scalable prefetcher training algorithm along with a hardware-based inter-thread prefetching mechanism. In some cases, blindly applying prefetching degrades performance. To reduce such negative effects, we propose an adaptive prefetch throttling scheme, which permits automatic GPGPU application- and hardware-specific adjustment. We show that adaptation reduces the negative effects of prefetching and can even improve performance. Overall, compared to the state-of-the-art software and hardware prefetching, our MT-prefetching improves performance on average by 16%(software pref.)/15% (hardware pref.) on our benchmarks.
  • Keywords
    computer graphic equipment; coprocessors; multi-threading; multiprocessing systems; storage management; CPU system; adaptive prefetch throttling; fine-grained thread; general-purpose GPU; hardware MT-prefetching; hardware-based interthread prefetching mechanism; many-thread aware prefetching mechanism; memory latency tolerance; multithreaded GPGPU; scalable prefetcher training algorithm; software MT-prefetching mechanism; thread-level parallelism; GPGPU; prefetch throttling; prefetching;
  • fLanguage
    English
  • Publisher
    ieee
  • Conference_Titel
    Microarchitecture (MICRO), 2010 43rd Annual IEEE/ACM International Symposium on
  • Conference_Location
    Atlanta, GA
  • ISSN
    1072-4451
  • Print_ISBN
    978-1-4244-9071-4
  • Type

    conf

  • DOI
    10.1109/MICRO.2010.44
  • Filename
    5695538