ENGLISH

Euro-Par 2022: Parallel Processing Workshops: Euro-Par 2022 International Workshops, Glasgow, UK, August 22–26, 2022, Revised Selected Papers

Book information

Publisher
Springer
Year
2023
ISBN
3031312082, 9783031312083
Language
english
Format
PDF
Filesize
20 MB (21336772 bytes)
Series
Lecture Notes in Computer Science, 13835
Pages
312\313
Time added
2023-05-08 01:20:07

Description

This book constitutes revised selected papers from the workshops held at the 28th International European Conference on Parallel and Distributed Computing, Euro-Par 2022, which took place in Glasgow, UK, in August 22–26, 2022 Out of a total of 35 submissions 24 papers have been accepted, 19 of these are included in this book. They stem from the following workshops: - Workshop on Algorithms, Models and Tools for Parallel Computing on Heterogeneous Platforms (HeteroPar)- Workshop on Asynchronous Many-Task systems for Exascale (AMTE) - Workshop on Domain Specific Languages for High-Performance Computing (DSL-HPC)- Workshop on Distributed and Heterogeneous Programming in C and C++ (DHPCC++)- Workshop on Resiliency in High Performance Computing in Clouds, Grids, and Clusters (Resilience) In addition, the proceedings also contains 6 extended abstracts from the PhD Symposium.  Preface Organization Contents AMTE Asynchronous Many-Task Systems for Exascale (AMTE) Workshop Description Organization Steering Committee Program Committee Quantifying Overheads in Charm++ and HPX Using Task Bench 1 Introduction 2 Related Work 3 Asynchronous Many-Task Systems 3.1 Charm++ 3.2 HPX 3.3 Commonalities and Differences 4 Task Bench 5 Improvements 5.1 Charm++ 5.2 HPX 6 Experiments 6.1 Performance of a Single Task on Each Core 6.2 Performance of Overdecomposition 6.3 Fine-grained Charm++ performance 7 Conclusion and Outlook References A Portable and Heterogeneous LU Factorization on IRIS 1 Introduction 2 IRIS Programming System 3 LU Factorization 4 Implementation 4.1 Memory Management 4.2 Tasking 4.3 GEMM 5 Performance Analysis 5.1 GETRF 5.2 TRSM-top 5.3 TRSM-Left 5.4 Join TRSM-Top and TRSM-Left 5.5 GEMM 5.6 Overall Performance 6 Related Works 7 Final Remarks and Future Directions References Halide Code Generation Framework in Phylanx 1 Introduction 2 Related Work 2.1 Domain Specific Languages (DSLs) 2.2 Code Generation Frameworks 2.3 AMT Runtime Systems 3 Enabling Technologies 3.1 HPX 3.2 Halide 3.3 APEX 3.4 Traveler 4 Performance Portability in Phylanx 4.1 HPX Runtime for Halide 4.2 Halide Integration in Phylanx 5 Results 5.1 System Setup 5.2 Experiments 6 Conclusion References DSL-HPC Workshop on Domain Specific Languages for High-Performance Computing (DSL-HPC) Workshop Description Organization Program Chair Program Committee Exploring the Suitability of the Cerebras Wafer Scale Engine for Stencil-Based Computation Codes*-4pt 1 Introduction 2 Background and Related Work 2.1 Cerebras Wafer Scale Engine 2.2 Programming the Wafer Scale Engine 3 TensorFlow for Encoding Stencil-Based Algorithms on the Wafer Scale Engine 4 Results 5 Conclusions References Performance of the Vipera Framework for DSLs on Micro-Core Architectures 1 Introduction 2 Related Work 2.1 Vipera Dynamic Language Framework 3 Experimental Environment 3.1 CPU Selection 3.2 LINPACK Overview 3.3 Sieve of Eratosthenes (Byte Sieve) Overview 4 Results and Discussion 4.1 LINPACK Runtime Performance 4.2 LINPACK Code Size 4.3 Sieve of Eratosthenes Runtime Performance 4.4 Sieve of Eratosthenes Code Size 4.5 Optimising Loops 5 Conclusion References FFTc: An MLIR Dialect for Developing HPC Fast Fourier Transform Libraries 1 Introduction 2 Background 3 Related Work 4 Methodology: A Domain-Specific Language for FFT 4.1 The FFTc Language and Grammar 4.2 FFTc Compilation Pipeline 5 Experimental Setup 6 Results 7 Discussion and Conclusion References Hetero-Par Workshop on Algorithms, Models and Tools for Parallel Computing on Heterogeneous Platforms (HeteroPar) Workshop Description Organization Steering Committee Program Chair Program Committee Programming Heterogeneous Architectures Using Hierarchical Tasks 1 Introduction 2 Related Work 3 Automatic Data Management 4 The Hierarchical Task Paradigm 4.1 Ensuring the Correctness of the DAG 5 Experimental Evaluation 6 Conclusion References A C++ Library for Memory Layout and Performance Portability of Scientific Applications 1 Introduction 2 From C++ Tuples to Compile-Time Data Structures 3 Generic Algorithms over Abstract Data Structures 4 Benchmarks 4.1 Memory Performance 4.2 Compute Performance 4.3 Application Example: Smoothed Particle Hydrodynamics 5 Conclusions References Implementation and Performance Evaluation of Memory System Using Addressable Cache for HPC Applications on HBM2 Equipped FPGAs 1 Introduction 2 Related Works 3 Proposed Memory System 3.1 Overview 3.2 System Design 3.3 Crossbar 3.4 Design of LocalStore 4 Design and Implementation of the API 4.1 Overview 4.2 RISC-V Core and Code Generation 5 Performance Evaluation 5.1 Environment and Program for Evaluation 5.2 Evaluation Result 6 Discussion 7 Conclusion and Future Work References Programming Abstractions for Preemptive Scheduling on FPGAs Using Partial Reconfiguration 1 Introduction 2 Related Work 3 The Controller Programming Model 4 Approach to Support Preemptive Scheduling on FPGAs 4.1 On-chip Infrastructure 4.2 Integration into the Controller Framework 4.3 Use Case: DPR Scheduler 5 Programmer's Abstractions 5.1 Kernel Interface Abstraction 5.2 Programmer Abstractions for Preemption 6 Experimental Study 6.1 Use Case: Scheduler of Randomly Generated Image Filter Tasks 6.2 Experimentation Environment 6.3 Results 7 Conclusions References Modeling Task Mapping for Data-Intensive Applications in Heterogeneous Systems 1 Introduction 2 State of the Art 3 Modeling 3.1 Abstract Model 3.2 Models for Different Design Stages 3.3 Extension: Full Usage of Data Busses 3.4 Extension: Streamability and Virtual Memory 4 Mixed-Integer Linear Programs 4.1 Device-Based ILP 4.2 Time-Based ILP 4.3 Extension: Streamable Devices 5 Evaluation 6 Conclusion References Mapping Tree-Shaped Workflows on Memory-Heterogeneous Architectures 1 Introduction 2 Related Work 3 Model 4 Heuristic Strategies 5 Experimental Evaluation 5.1 Experimental Setup 5.2 Results 6 Conclusions and Future Work References Hetero-Vis: A Framework for Latency Optimized Heterogeneous Deployment of Convolutional Neural Networks 1 Introduction 2 Preliminaries 3 Related Work 4 Hetero-Vis Framework 4.1 Methodology 4.2 Framework Components 4.3 Algorithm for Hetero-Vis Framework: 5 Results and Discussions 5.1 Experimental Setup 5.2 Experiment Outcomes 6 Limitations and Future Work 7 Conclusion References Rapid Development of OS Support with PMCSched for Scheduling on Asymmetric Multicore Systems 1 Introduction 2 Related Work 3 PMCSched: Implementation Challenges and Design 4 Experimental Case Study 5 Conclusions and Future Work References HIPLZ: Enabling Performance Portability for Exascale Systems 1 Introduction 2 Background 2.1 Heterogeneous-compute Interface for Portability (HIP) 2.2 Standard Portable Intermediate Representation (SPIR-V) and Fat Binary 2.3 OpenCL and HIPCL 2.4 Level Zero Runtime 3 Design and Implementation 3.1 Design Goal 3.2 The Compilation System 3.3 Runtime System 3.4 Streams 3.5 Memory Management 3.6 Kernel and Module Management 3.7 Device Management 3.8 SYCL Inter-operation 3.9 Kernel Library 3.10 Discussion 4 Evaluation 4.1 Employed GPU System 4.2 Overview of Tests 4.3 Results 5 Related Work 6 Conclusion References StorAlloc: A Simulator for Job Scheduling on Heterogeneous Storage Resources 1 Introduction 2 Context and Motivation 3 Related Work 4 Architecture 4.1 Scheduling of Storage Requests 4.2 Storage Abstraction 4.3 Simulation 4.4 Implementation Details 5 Evaluation 5.1 Simulation Setup 5.2 Analysis 6 Conclusion References Performance and Scalability Analysis of AI-Accelerated CFD Simulations Across Various Computing Platforms 1 Introduction 2 Related Work 3 AI-Based Acceleration of CFD Simulations for Chemical Mixing Using CFD Suite 4 Methodology of Benchmarking and Analysis 4.1 Hardware and Software Environments 4.2 Benchmarking Scenarios 5 Performance and Scalability Analysis: Training 6 Performance Analysis: Inference 7 Conclusions References IV Misc Miscellaneous Workshops Workshop Description Organization Workshop Chairs Workshop Reviewers Performance Portability Assessment: Non-negative Matrix Factorization as a Case Study 1 Introduction 2 Background and Related Work 3 Non-negative Factorization 4 NMF Code Implementation 4.1 BLAS Baseline Implementation 4.2 SYCL Implementation 4.3 OpenMP Implementation 4.4 OpenMP and SYCL Common Ground 5 Experimental Conditions 5.1 Work Environment 5.2 Data Description 5.3 Other Considerations 6 Discussion 6.1 CPU Discussion 6.2 GPU Discussion 7 Conclusion References Task-Level Checkpointing System for Task-Based Parallel Workflows 1 Introduction 2 Checkpointing Task-Based Workflows 3 Solution Design and Implementation 4 Evaluation 4.1 Checkpointing Overhead 4.2 Recovery Speedup 4.3 Avoid Checkpointing Tasks 4.4 Customized Policies 5 Related Work 6 Conclusion References PhD Symposium Euro-Par PhD Symposium PhD Symposium Description Organization Chair Program Committee A Stochastic Programming Approach for an Enhanced Performance of a Multi-committees Byzantine Fault Tolerant Algorithm 1 Introduction 2 Stochastic Model 2.1 System Assumption 2.2 Decision Variables and Objective Function 2.3 Performance Constraints 3 Performance Evaluation 4 Conclusion References Coupe: A Modular, Multi-threaded Mesh Partitioning Platform 1 Introduction 2 The Mesh Partitioning Problem 2.1 Topologic vs. Geometrical Mesh Partitioning 2.2 Current Partitioning Platforms and Features 3 Coupe: A Dedicated Mesh Partitioning Platform 3.1 Software Choices and Architecture 3.2 Example of Scaling Algorithm Under Development 4 Outlook References Preliminary Study of Resource Allocation in Wireless Communications 1 Introduction 2 System Model and Existing Work 2.1 Sum SE Maximisation Under Power Constraints 2.2 Sum SE Maximisation Without Power Considerations 3 Our Contributions 4 Conclusion and Future Directions References Benchmarking Parallelism in Unikernels 1 Introduction 1.1 Unikernels for the Cloud 1.2 Parallel Performance of Unikernels 2 Experimental Design 2.1 Benchmark Metrics 2.2 The Mandelbrot Benchmark 2.3 Comparators 3 Results 3.1 Wall Clock Run Times 3.2 Boot up Times 3.3 Parallel Speedups 3.4 Discussion 4 Conclusion References Machine Learning Methodologies to Support HPC Systems Operations: Anomaly Detection 1 Introduction 1.1 Contributions 1.2 Anomalies and Dataset 2 The LSTM Autoencoder Network 3 Experimental Results 4 Conclusions References FPGAs in Supercomputers: Performance Through Dataflow Programming and Flexibility 1 Research Problem 1.1 Dataflow Programming 1.2 Flexibility 2 Dataflow: Compiler Technology and DSLs 2.1 Vitis HLS Open Source Front-End 2.2 An Example: The Stencil Pragma 2.3 DSLs and xDSL 3 Flexibility: Dynamic Partial Reconfiguration in Task-Based Models References Author Index

Similar books