• DocumentCode
    1882594
  • Title

    A Mostly-Clean DRAM Cache for Effective Hit Speculation and Self-Balancing Dispatch

  • Author

    Jaewoong Sim ; Loh, Gabriel H. ; Hyesoon Kim ; O´Connor, M. ; Thottethodi, Mithuna

  • fYear
    2012
  • fDate
    1-5 Dec. 2012
  • Firstpage
    247
  • Lastpage
    257
  • Abstract
    Die-stacking technology allows conventional DRAM to be integrated with processors. While numerous opportunities to make use of such stacked DRAM exist, one promising way is to use it as a large cache. Although previous studies show that DRAM caches can deliver performance benefits, there remain inefficiencies as well as significant hardware costs for auxiliary structures. This paper presents two innovations that exploit the bursty nature of memory requests to streamline the DRAM cache. The first is a low-cost Hit-Miss Predictor (HMP) that virtually eliminates the hardware overhead of the previously proposed multi-megabyte Miss Map structure. The second is a Self-Balancing Dispatch (SBD) mechanism that dynamically sends some requests to the off-chip memory even though the request may have hit in the die-stacked DRAM cache. This makes effective use of otherwise idle off-chip bandwidth when the DRAM cache is servicing a burst of cache hits. These techniques, however, are hampered by dirty (modified) data in the DRAM cache. To ensure correctness in the presence of dirty data in the cache, the HMP must verify that a block predicted as a miss is not actually present, otherwise the dirty block must be provided. This verification process can add latency, especially when DRAM cache banks are busy. In a similar vein, SBD cannot redirect requests to off-chip memory when a dirty copy of the block exists in the DRAM cache. To relax these constraints, we introduce a hybrid write policy for the cache that simultaneously supports write-through and write-back policies for different pages. Only a limited number of pages are permitted to operate in a write-back mode at one time, thereby bounding the amount of dirty data in the DRAM cache. By keeping the majority of the DRAM cache clean, most HMP predictions do not need to be verified, and the self balancing dispatch has more opportunities to redistribute requests (i.e., only requests to the limited number of dirty pages must go to th- DRAM cache to maintain correctness). Our proposed techniques improve performance compared to the Miss Map-based DRAM cache approach while simultaneously eliminating the costly Miss Map structure.
  • Keywords
    DRAM chips; cache storage; HMP predictions; SBD mechanism; auxiliary structures; cache hits; die-stacked DRAM cache; die-stacking technology; dirty data; effective hit speculation; hardware costs; hardware overhead elimination; hybrid write policy; low-cost hit-miss predictor; multimegabyte MissMap structure; off-chip bandwidth; off-chip memory; performance benefits; self-balancing dispatch; verification process; write-back policies; write-through policies;
  • fLanguage
    English
  • Publisher
    ieee
  • Conference_Titel
    Microarchitecture (MICRO), 2012 45th Annual IEEE/ACM International Symposium on
  • Conference_Location
    Vancouver, BC
  • ISSN
    1072-4451
  • Print_ISBN
    978-1-4673-4819-5
  • Type

    conf

  • DOI
    10.1109/MICRO.2012.31
  • Filename
    6493624