Reducing cache and TLB power by exploiting memory region and privilege level semantics

被引：1

作者：

Fang, Zhen ^{[1
]}

Zhao, Li ^{[2
]}

Jiang, Xiaowei ^{[2
]}

Lu, Shih-lien ^{[2
]}

Iyer, Ravi ^{[2
]}

Li, Tong ^{[2
]}

Lee, Seung Eun ^{[3
]}

机构：

[1] AMD Corp, Austin, TX 78735 USA

[2] Intel Corp, Hillsboro, OR 97124 USA

[3] Seoul Natl Univ Sci & Technol, Elect & Informat Engn Dept, Seoul, South Korea

来源：

JOURNAL OF SYSTEMS ARCHITECTURE | 2013年 / 59卷 / 06期

关键词：

First-level cache; Translation lookaside buffer; Memory regions; Ring level; Simulation; BEHAVIOR;

D O I：

10.1016/j.sysarc.2013.04.002

中图分类号：

TP3 [计算技术、计算机技术];

学科分类号：

0812 ;

摘要：

The L1 cache in today's high-performance processors accesses all ways of a selected set in parallel. This constitutes a major source of energy inefficiency: at most one of the N fetched blocks can be useful in an N-way set-associative cache. The other N-1 cachelines will all be tag mismatches and subsequently discarded. We propose to eliminate unnecessary associative fetches by exploiting certain software semantics in cache design, thus reducing dynamic power consumption. Specifically, we use memory region information to eliminate unnecessary fetches in the data cache, and ring level information to optimize fetches in the instruction cache. We present a design that is performance-neutral, transparent to applications, and incurs a space overhead of mere 0.41% of the L1 cache. We show significantly reduced cache lookups with benchmarks including SPEC CPU, SPECjbb, SPECjApp-Server, PARSEC, and Apache. For example, for SPEC CPU 2006, the proposed mechanism helps to reduce cache block fetches from the data and instruction caches by an average of 29% and 53% respectively, resulting in power savings of 17% and 35% in the caches, compared to the aggressively clock-gated baselines. (C) 2013 Elsevier B.V. All rights reserved.

引用

页码：279 / 295

页数：17

共 6 条

[1] Reducing tag activities for power efficiency in I-cache memory
Zhu Xiaoping
Tiow, Tay Teng
2006 INTERNATIONAL CONFERENCE ON COMMUNICATIONS, CIRCUITS AND SYSTEMS PROCEEDINGS, VOLS 1-4: VOL 1: SIGNAL PROCESSING, 2006, : 2766 - 2770
[2] Exploiting Flash Memory for Reducing Disk Power Consumption in Portable Media Players
Kim, Jaewoo
Yang, Ahron
Song, Minseok
IEEE TRANSACTIONS ON CONSUMER ELECTRONICS, 2009, 55 (04) : 1997 - 2004
[3] Exploiting OS-Level Memory Offlining for DRAM Power Management
Lee, Seunghak
Kim, Nam Sung
Kim, Daehoon
IEEE COMPUTER ARCHITECTURE LETTERS, 2019, 18 (02) : 141 - 144
[4] Lowering Latency of Embedded Memory by Exploiting In-Cell Victim Cache Hierarchy Based on Emerging Multi- Level Memory Devices
Wu, Juejian
Liao, Tianyu
Li, Taixin
Xu, Yixin
Narayanan, Vijaykrishnan
Liu, Yongpan
Yang, Huazhong
Li, Xueqing
2023 IEEE/ACM INTERNATIONAL CONFERENCE ON COMPUTER AIDED DESIGN, ICCAD, 2023,
[5] EXTREME: Exploiting Page Table for Reducing Refresh Power of 3D-Stacked DRAM Memory
Shin, Ho Hyun
Park, Young Min
Choi, Duheon
Kim, Byoung Jin
Cho, Dae-Hyung
Chung, Eui-Young
IEEE TRANSACTIONS ON COMPUTERS, 2018, 67 (01) : 32 - 44
[6] A low-power 2.5-GHz 90-nm level 1 cache and memory management unit
Haigh, JR
Wilkerson, MW
Miller, JB
Beatty, TS
Strazdus, SJ
Clark, LT
IEEE JOURNAL OF SOLID-STATE CIRCUITS, 2005, 40 (05) : 1190 - 1199

← 1 →