Successive Refinement in Large-Scale Computation: Expediting Model Inference Applications

被引：0

作者：

Esfahanizadeh, Homa ^{[1
]}

Cohen, Alejandro ^{[2
]}

Shamai, Shlomo ^{[2
]}

Medard, Muriel ^{[3
]}

机构：

[1] Nokia Bell Labs, Murray Hill, NJ 07974 USA

[2] Technion Israel Inst Technol, Elect & Comp Engn Dept, IL-3200003 Haifa, Israel

[3] MIT, Res Lab Elect RLE, Cambridge, MA 02139 USA

来源：

IEEE TRANSACTIONS ON SIGNAL PROCESSING | 2025年 / 73卷

基金：

美国国家科学基金会; 以色列科学基金会;

关键词：

Computation; layered resolution; adaptive; linear; nonlinear; machine learning; inference; matrix multiplication;

D O I：

10.1109/TSP.2025.3537409

中图分类号：

TM [电工技术]; TN [电子技术、通信技术];

学科分类号：

0808 ; 0809 ;

摘要：

Modern computationally-intensive applications often operate under time constraints, necessitating acceleration methods and distribution of computational workloads across multiple entities. However, the outcome is either achieved within the desired timeline or not, and in the latter case, valuable resources are wasted. In this paper, we introduce solutions for layered-resolution computation. These solutions allow lower-resolution results to be obtained at an earlier stage than the final result. This innovation notably enhances the deadline-based systems, as if a computational job is terminated due to time constraints, an approximate version of the final result can still be generated. Moreover, in certain operational regimes, a high-resolution result might be unnecessary, because the low-resolution result may already deviate significantly from the decision threshold, for example in AI-based decision-making systems. Therefore, operators can decide whether higher resolution is needed or not based on intermediate results, enabling computations with adaptive resolution. We present our framework for two critical and computationally demanding jobs: distributed matrix multiplication (linear) and model inference in machine learning (nonlinear). Our theoretical and empirical results demonstrate that the execution delay for the first resolution is significantly shorter than that for the final resolution, while maintaining overall complexity comparable to the conventional one-shot approach. Our experiments further illustrate how the layering feature increases the likelihood of meeting deadlines and enables adaptability and transparency in massive, large-scale computations.

引用

页码：811 / 826

页数：16

共 50 条

[31] SUPERCONDUCTIVITY - LARGE-SCALE APPLICATIONS
HEIN, RA
SCIENCE, 1974, 185 (4147) : 211 - 222
[32] LARGE-SCALE APPLICATIONS OF SUPERCONDUCTIVITY
FONER, S
SCHWARTZ, BB
JOURNAL OF THE ELECTROCHEMICAL SOCIETY, 1979, 126 (03) : C153 - C153
[33] Some large-scale matrix computation problems
Bai, ZJ
Fahey, M
Golub, G
JOURNAL OF COMPUTATIONAL AND APPLIED MATHEMATICS, 1996, 74 (1-2) : 71 - 89
[34] LARGE-SCALE SCIENTIFIC COMPUTATION VIA MINICOMPUTER
SCHAEFER, HF
MILLER, WH
COMPUTERS & CHEMISTRY, 1977, 1 (02): : 85 - 90
[35] Large-scale FDTD computation as computational electromagnetics
Kashiwa, Tatsuya
IEEJ Transactions on Fundamentals and Materials, 2009, 129 (02) : 50 - 53
[36] EFFICIENT MEMORY ACCESS IN LARGE-SCALE COMPUTATION
VITTER, JS
LECTURE NOTES IN COMPUTER SCIENCE, 1991, 480 : 26 - 41
[37] Large-scale logic-array computation
Margolus, N
HIGH-SPEED COMPUTING, DIGITAL SIGNAL PROCESSING, AND FILTERING USING RECONFIGURABLE LOGIC, 1996, 2914 : 341 - 352
[38] Quantum computation for large-scale image classification
Ruan, Yue
Chen, Hanwu
Tan, Jianing
Li, Xi
QUANTUM INFORMATION PROCESSING, 2016, 15 (10) : 4049 - 4069
[39] Large-scale computation at sharc-net
Couchman, H
HIGH PERFORMANCE COMPUTING SYSTEMS AND APPLICATIONS, 2003, 727 : 33 - 33
[40] LARGE-SCALE COMPUTATION FOR MODELING PATIENTS ON A MICROCOMPUTER
DELAND, EC
KUN, LG
IEEE TRANSACTIONS ON BIOMEDICAL ENGINEERING, 1984, 31 (08) : 569 - 569

← 1 2 3 4 5 →