DocumentCode
2582221
Title
CAPRI: Prediction of compaction-adequacy for handling control-divergence in GPGPU architectures
Author
Rhu, Minsoo ; Erez, Mattan
Author_Institution
Electr. & Comput. Eng. Dept., Univ. of Texas at Austin, Austin, TX, USA
fYear
2012
fDate
9-13 June 2012
Firstpage
61
Lastpage
71
Abstract
Wide SIMD-based GPUs have evolved into a promising platform for running general purpose workloads. Current programmable GPUs allow even code with irregular control to execute well on their SIMD pipelines. To do this, each SIMD lane is considered to execute a logical thread where hardware ensures that control flow is accurate by automatically applying masked execution. The masked execution, however, often degrades performance because the issue slots of masked lanes are wasted. This degradation can be mitigated by dynamically compacting multiple unmasked threads into a single SIMD unit. This paper proposes a fundamentally new approach to branch compaction that avoids the unnecessary synchronization required by previous techniques and that only stalls threads that are likely to benefit from compaction. Our technique is based on the compaction-adequacy predictor (CAPRI). CAPRI dynamically identifies the compaction-effectiveness of a branch and only stalls threads that are predicted to benefit from compaction. We utilize a simple single-level branch-predictor inspired structure and show that this simple configuration attains a prediction accuracy of 99.8% and 86.6% for non-divergent and divergent workloads, respectively. Our performance evaluation demonstrates that CAPRI consistently outperforms both the baseline design that never attempts compaction and prior work that stalls upon all divergent branches.
Keywords
graphics processing units; parallel architectures; CAPRI; GPGPU Architectures; SIMD pipelines; SIMD-based GPU; branch compaction; compaction-adequacy prediction; control-divergence; divergent workloads; general purpose workloads; logical thread; nondivergent workloads; unmasked threads; Compaction; Computer architecture; Graphics processing unit; Hardware; Instruction sets; Synchronization; Vectors;
fLanguage
English
Publisher
ieee
Conference_Titel
Computer Architecture (ISCA), 2012 39th Annual International Symposium on
Conference_Location
Portland, OR
ISSN
1063-6897
Print_ISBN
978-1-4673-0475-7
Electronic_ISBN
1063-6897
Type
conf
DOI
10.1109/ISCA.2012.6237006
Filename
6237006
Link To Document