• DocumentCode
    2582221
  • Title

    CAPRI: Prediction of compaction-adequacy for handling control-divergence in GPGPU architectures

  • Author

    Rhu, Minsoo ; Erez, Mattan

  • Author_Institution
    Electr. & Comput. Eng. Dept., Univ. of Texas at Austin, Austin, TX, USA
  • fYear
    2012
  • fDate
    9-13 June 2012
  • Firstpage
    61
  • Lastpage
    71
  • Abstract
    Wide SIMD-based GPUs have evolved into a promising platform for running general purpose workloads. Current programmable GPUs allow even code with irregular control to execute well on their SIMD pipelines. To do this, each SIMD lane is considered to execute a logical thread where hardware ensures that control flow is accurate by automatically applying masked execution. The masked execution, however, often degrades performance because the issue slots of masked lanes are wasted. This degradation can be mitigated by dynamically compacting multiple unmasked threads into a single SIMD unit. This paper proposes a fundamentally new approach to branch compaction that avoids the unnecessary synchronization required by previous techniques and that only stalls threads that are likely to benefit from compaction. Our technique is based on the compaction-adequacy predictor (CAPRI). CAPRI dynamically identifies the compaction-effectiveness of a branch and only stalls threads that are predicted to benefit from compaction. We utilize a simple single-level branch-predictor inspired structure and show that this simple configuration attains a prediction accuracy of 99.8% and 86.6% for non-divergent and divergent workloads, respectively. Our performance evaluation demonstrates that CAPRI consistently outperforms both the baseline design that never attempts compaction and prior work that stalls upon all divergent branches.
  • Keywords
    graphics processing units; parallel architectures; CAPRI; GPGPU Architectures; SIMD pipelines; SIMD-based GPU; branch compaction; compaction-adequacy prediction; control-divergence; divergent workloads; general purpose workloads; logical thread; nondivergent workloads; unmasked threads; Compaction; Computer architecture; Graphics processing unit; Hardware; Instruction sets; Synchronization; Vectors;
  • fLanguage
    English
  • Publisher
    ieee
  • Conference_Titel
    Computer Architecture (ISCA), 2012 39th Annual International Symposium on
  • Conference_Location
    Portland, OR
  • ISSN
    1063-6897
  • Print_ISBN
    978-1-4673-0475-7
  • Electronic_ISBN
    1063-6897
  • Type

    conf

  • DOI
    10.1109/ISCA.2012.6237006
  • Filename
    6237006