DocumentCode
1009209
Title
High Performance, Energy Efficiency, and Scalability With GALS Chip Multiprocessors
Author
Yu, Zhiyi ; Baas, Bevan M.
Author_Institution
Dept. of Microelectron., Fudan Univ., Shanghai
Volume
17
Issue
1
fYear
2009
Firstpage
66
Lastpage
79
Abstract
Chip multiprocessors with globally asynchronous locally synchronous (GALS) clocking styles are promising candidates for processing computationally-intensive and energy-constrained workloads. The GALS methodology simplifies clock tree design, provides opportunities to use clock and voltage scaling jointly in system submodules to achieve high energy efficiencies, and can also result in easily scalable clocking systems. However, its use typically also introduces performance penalties due to additional communication latency between clock domains. We show that GALS chip multiprocessors (CMPs) with large inter-processor first-inputs-first-outputs (FIFOs) buffers can inherently hide much of the GALS performance penalty while executing applications that have been mapped with few communication loops. In fact, the penalty can be driven to zero with sufficiently large FIFOs and the removal of multiple-loop communication links. We present an example mesh-connected GALS chip multiprocessor and show it has a less than 1% performance (throughput) reduction on average compared to the corresponding synchronous system for many DSP workloads. Furthermore, adaptive clock and voltage scaling for each processor provides an approximately 40% power savings without any performance reduction. These results compare favorably with the GALS uniprocessor, which compared to the corresponding synchronous uniprocessor, has a reported greater than 10% performance (throughput) reduction and an energy savings of approximately 25% using dynamic clock and voltage scaling for many general purpose applications.
Keywords
asynchronous circuits; clocks; integrated circuit design; microprocessor chips; DSP workload; FIFO buffer; GALS methodology; GALS uniprocessor comparison; clock tree design; computationally-intensive workloads; energy-constrained workloads; globally asynchronous locally synchronous clocking styles; inter-processor first-inputs-first-outputs; mesh-connected GALS chip multiprocessor; multipleloop communication links; scalable clocking systems; system submodules; voltage scaling; Array processor; chip multiprocessor; energy efficient; globally asynchronous locally synchronous (GALS); low power; scalable;
fLanguage
English
Journal_Title
Very Large Scale Integration (VLSI) Systems, IEEE Transactions on
Publisher
ieee
ISSN
1063-8210
Type
jour
DOI
10.1109/TVLSI.2008.2001947
Filename
4689315
Link To Document