Large-Scale Experiment for Topology-Aware Resource Management

被引：0

作者：

Georgiou, Yiannis ^{[1
]}

Mercier, Guillaume ^{[2
]}

Villiermet, Adele ^{[3
]}

机构：

[1] Atos Bull, Grenoble, France

[2] Bordeaux INP, Talence, France

[3] Inria Bordeaux Sud Ouest, Talence, France

来源：

EURO-PAR 2017: PARALLEL PROCESSING WORKSHOPS | 2018年 / 10659卷

关键词：

Resource management; Job allocation; Topology-aware placement; Scheduling; SLURM; PLACEMENT;

D O I：

10.1007/978-3-319-75178-8_15

中图分类号：

TP301 [理论、方法];

学科分类号：

081202 ;

摘要：

A Resource and Job Management System (RJMS) is a crucial system software part of the HPC stack. It is responsible for efficiently delivering computing power to applications in supercomputing environments and its main intelligence relies on resource selection techniques to find the most adapted resources to schedule the users' jobs. In [8], we introduced a new topology-aware resource selection algorithm to determine the best choice among the available nodes of the platform based on their position in the network and on application behaviour (expressed as a communication matrix). We did integrate this algorithm as a plugin in SLURM and validated it with several optimization schemes by making comparisons with the default SLURM algorithm. This paper presents further experiments with regard to this selection process.

引用

页码：179 / 186

页数：8

共 50 条

[1] Topology-aware algorithms for large-scale communication
Rodrigues, L
Veríssimo, P
ADVANCES IN DISTRIBUTED SYSTEMS: ADVANCED DISTRIBUTED COMPUTING: FROM ALGORITHMS TO SYSTEMS, 2000, 1752 : 127 - 156
[2] Topology-aware algorithms for large-scale communication
Rodrigues, Luís
Veríssimo, Paulo
Lecture Notes in Computer Science (including subseries Lecture Notes in Artificial Intelligence and Lecture Notes in Bioinformatics), 2000, 1752 : 127 - 156
[3] Topology-Aware Mappings for Large-Scale Eigenvalue Problems
Aktulga, Hasan Metin
Yang, Chao
Ng, Esmond G.
Maris, Pieter
Vary, James P.
EURO-PAR 2012 PARALLEL PROCESSING, 2012, 7484 : 830 - 842
[4] Topology-aware Sparse Allreduce for Large-scale Deep Learning
Thao Nguyen Truong
Wahib, Mohamed
Takano, Ryousei
2019 IEEE 38TH INTERNATIONAL PERFORMANCE COMPUTING AND COMMUNICATIONS CONFERENCE (IPCCC), 2019,
[5] A scheduling framework for large-scale, parallel, and topology-aware applications
Kravtsov, Valentin
Bar, Pavel
Carmeli, David
Schuster, Assaf
Swain, Martin
JOURNAL OF PARALLEL AND DISTRIBUTED COMPUTING, 2010, 70 (09) : 983 - 992
[6] Topology-aware resource management for HPC applications
Georgiou, Yiannis
Jeannot, Emmanuel
Mercier, Guillaume
Villiermet, Adele
18TH INTERNATIONAL CONFERENCE ON DISTRIBUTED COMPUTING AND NETWORKING (ICDCN 2017), 2017,
[7] Topology-Aware Data Aggregation for Intensive I/O on Large-Scale Supercomputers
Tessier, Francois
Malakar, Preeti
Vishwanath, Venkatram
Jeannot, Emmanuel
Isaila, Florin
PROCEEDINGS OF FIRST WORKSHOP ON OPTIMIZATION OF COMMUNICATION IN HPC RUNTIME SYSTEMS (COM-HPC 2016), 2016, : 73 - 81
[8] Topology-aware Virtual Resource Management for Heterogeneous Multicore Systems
Qian, Jianmin
Li, Jian
Ma, Ruhui
PROCEEDINGS OF THE 2018 DESIGN, AUTOMATION & TEST IN EUROPE CONFERENCE & EXHIBITION (DATE), 2018, : 177 - 182
[9] TAPIOCA: An I/O Library for Optimized Topology-Aware Data Aggregation on Large-Scale Supercomputers
Tessier, Francois
Vishwanath, Venkatram
Jeannot, Emmanuel
2017 IEEE INTERNATIONAL CONFERENCE ON CLUSTER COMPUTING (CLUSTER), 2017, : 70 - 80
[10] Improving Large-scale Storage System Performance via Topology-aware and Balanced Data Placement
Wang, Feiyi
Oral, Sarp
Gupta, Saurabh
Tiwari, Devesh
Vazhkudai, Sudharshan S.
2014 20TH IEEE INTERNATIONAL CONFERENCE ON PARALLEL AND DISTRIBUTED SYSTEMS (ICPADS), 2014, : 656 - 663

← 1 2 3 4 5 →