Network and Parallel Computing: 19th IFIP WG 10.3 International Conference, NPC 2022, Jinan, China, September 24–25, 2022, Proceedings
Book information
Description
This book constitutes the proceedings of the 19th IFIP WG 10.3 International Conference on Network and Parallel Computing, NPC 2022, which was held in Jinan, China, during September 24-25, 2022. The 23 full papers and 8 short papers presented in this volume were carefully reviewed and selected from 89 submissions. They were organized in topical sections as follows: computer architecture; cloud computing; deep learning; emerging applications; and storage and IO. Preface Organization Contents Architecture A Routing-Aware Mapping Method for Dataflow Architectures 1 Introduction 2 Background and Related Works 2.1 Dataflow Architecture 2.2 Related Works 3 Motivation 4 Our Method 5 Evaluation 5.1 Methodology 5.2 Performance Improvement 5.3 Energy Saving 5.4 Scalability 5.5 Compilation Time 6 Conclusion References Optimizing Winograd Convolution on GPUs via Partial Kernel Fusion 1 Introduction 2 Background 2.1 Implementations of Convolution 2.2 Winograd Convolution 2.3 NVIDIA GPU Architecture and Tensor Cores 3 Related Work 4 Methodology 4.1 Optimizing EWMM Stage 4.2 PKF (Partial Kernel Fusion) 5 Implementation and Experiment 5.1 Implementation PKF on TVM 5.2 Experiment 6 Conclusion References Adaptive Low-Cost Loop Expansion for Modulo Scheduling 1 Introduction 2 Expanded Modulo Scheduling 2.1 Data Dependence Graph 2.2 Expansion Count and Iteration Interval 2.3 Scheduling 2.4 Resolving Expansion Faults 2.5 Completing the MRT 3 Performance Evaluation 3.1 Target Architecture 3.2 Adaptation of EMS 3.3 Experiment Setup 3.4 Experiment Results 4 Conclusion References SADD: A Novel Systolic Array Accelerator with Dynamic Dataflow for Sparse GEMM in Deep Learning 1 Introduction 2 Background 2.1 Dataflows in the Systolic Array 2.2 Sparsity 3 SADD Architecture 3.1 Group-Structure-Maintained Compression 3.2 The SIS and SWS 3.3 The Performance of SIS and SWS with Different GEMM Sizes 3.4 The SDD and SADD 4 Experimental Results 4.1 Experimental Setup 4.2 Performance Comparison of Different Dataflows 4.3 Comparison of the SADD and the TPU 4.4 Scalability Analysis 4.5 Hardware Cost Analysis 5 Related Work 6 Conclusion References CSR&RV: An Efficient Value Compression Format for Sparse Matrix-Vector Multiplication 1 Introduction 2 The Compressed Sparse Row and Repetition Value Format 2.1 CSR&RV Representation 2.2 SpMV Algorithm 3 Experimental Results 3.1 Performance Comparison 3.2 Memory Overhead 3.3 Pre-processing 4 Conclusion References Rgs-SpMM: Accelerate Sparse Matrix-Matrix Multiplication by Row Group Splitting Strategy on the GPU 1 Introduction 2 Related Work and Motivation 3 Rgs-SpMM Design 3.1 Data Organization in Rgs-SpMM 3.2 Row Group Splitting 4 Experiment Evaluations 4.1 Overall Performance 4.2 Analysis of Results 5 Conclusion References Cloud Computing Interference-aware Workload Scheduling in Co-located Data Centers 1 Introduction 2 Related Work 3 Interference-aware Solution 3.1 Performance Interference Metric 3.2 Performance Interference Model Based on Linear Regression 3.3 Interference-aware Workload Scheduling 4 Experiment and Evaluation 4.1 Prediction Accuracy of Performance Interference Models 4.2 Evaluation of Scheduling Strategies on Throughput 5 Conclusion References FaaSPipe: Fast Serverless Workflows on Distributed Shared Memory 1 Introduction 2 Related Work 3 Design and Implementation 3.1 The PipeFunc Programming Model 3.2 System Architechture of FaaSPipe 3.3 Intra-workflow Memory Sharing 3.4 Full-Duplex Memory Transfer 4 Evaluation 4.1 Distributed Word Count 4.2 LightGBM 4.3 Efficency of FaaSPipe vs. Faasm 5 Conclusion References TopKmer: Parallel High Frequency K-mer Counting on Distributed Memory 1 Introduction 2 Background 2.1 Parallel K-mer Counting 2.2 Heavy Hitters 3 TopKmer Counter 3.1 Multi-layer Hash Table 3.2 Insert 3.3 Query 4 Parallel K-Mer Counting Framework 5 Results 5.1 Experiment Setup 5.2 Quality of Counting 5.3 Performance Comparison 5.4 Scaling Capability 5.5 Time Consumption Analysis 6 Conclusion References Flexible Supervision System: A Fast Fault-Tolerance Strategy for Cloud Applications in Cloud-Edge Collaborative Environments 1 Introduction 2 Flexible Supervision System Architecture 3 Fault Detection and Fault-Tolerance Strategy 4 Experimental Evaluation 5 Related Work 6 Conclusions and Future Work References Adjust: An Online Resource Adjustment Framework for Microservice Programs 1 Introduction 2 A QoS Awareness Framework for Microservices 2.1 Microservice Analyzer (MSA) 2.2 Microservice Prediction Model (MSPM) 2.3 Microservice Performance Guarantor (MSPG) 3 Evaluation 3.1 Performance Guarantee 3.2 Resource Re-collection 4 Conclusion References Cloud-Native Server Consolidation for Energy-Efficient FaaS Deployment 1 Introduction 2 Key Design Considerations 3 DAC Design 3.1 System Overview 3.2 Function Classifier 3.3 Consolidation Controller 4 Evaluation 4.1 Methodologies 4.2 Evaluation Results 5 Conclusion References Deep Learning NeuProMa: A Toolchain for Mapping Large-Scale Spiking Convolutional Neural Networks onto Neuromorphic Processor 1 Introduction 2 Background 2.1 Neuromorphic Processor 2.2 Spiking Convolutional Neural Network 3 Related Work 4 NeuProMa 4.1 Splitting 4.2 Partitioning 4.3 Mapping 5 Experiment Setup 5.1 Experiment Platform 5.2 Evaluated SCNNs 6 Experiment Results 6.1 Splitting Performance 6.2 Partitioning and Mapping Performance 7 Conclusion References Multi-clusters: An Efficient Design Paradigm of NN Accelerator Architecture Based on FPGA 1 Introduction 2 Background 2.1 Design Patterns of Accelerator 2.2 Related Work 3 Overall Method 3.1 Division Method 3.2 Architecture Design 3.3 Design Space Exploration 3.4 Scheduling Strategy 4 Experiment 4.1 Experiment Setup 4.2 Comparison with CPU and GPU 4.3 Comparison with Previous FPGA Accelerators 5 Conclusion References TrainFlow: A Lightweight, Programmable ML Training Framework via Serverless Paradigm 1 Introduction 2 Background and Challenges 2.1 Distributed ML Training 2.2 Challenges 3 TrainFlow Design 3.1 Overview 3.2 Serverless Process Model and Training Basics 3.3 Programmability Extension with Event-Driven Hook 4 Implementation 5 Evaluation 5.1 Availability 5.2 Programmability 6 Related Work 6.1 Classic ML Training 6.2 Serverless ML Training 7 Conclusion References DRP:Discrete Rank Pruning for Neural Network 1 Introduction 2 Related Work 2.1 Compression Techniques 2.2 Sparse Method 2.3 Structured Pruning 3 Consideration Bias Sparsity 4 Discrete Rank Pruning 5 Experiment and Evaluation 5.1 Datasets and Network Models 5.2 Implementation 5.3 Results and Analysis on CBS 5.4 Results and Analysis on DRP 6 Conclusion References TransMigrator: A Transformer-Based Predictive Page Migration Mechanism for Heterogeneous Memory 1 Introduction 2 TransMigrator 2.1 Design of Neural Network 2.2 Page Migration 3 Evaluation and Analysis 3.1 Trace Collection 3.2 Network Training 3.3 Migration Simulation 3.4 Access Time 3.5 Energy Consumption 3.6 Network Overhead 4 Related Work 5 Conclusion References Hardware Acceleration for 1D-CNN Based Real-Time Edge Computing 1 Introduction 2 Background 2.1 State-of-the-Art CNN Accelerators 2.2 CNN in Real-Time Computing 3 Proposed Architecture for 1D-CNN 3.1 Data Reuse 3.2 Accelerated 1D-CNN Architecture 3.3 Compiler for 1D-CNN Architecture Generation 4 Results 4.1 Setup 4.2 Evaluations of Power, Latency and Bandwidth 4.3 Comparative Analysis 5 Conclusion References Emerging Applications DLC: An Optimization Framework for Full-State Quantum Simulation 1 Introduction 2 Background and Related Work 2.1 Quantum States and Quantum Circuits 2.2 Full-State Quantum Simulator 2.3 Related Work 3 Framework Overview 4 CPU-GPU Locality Enhancement 4.1 Data Dependency Analysis 4.2 CPU-GPU Locality Enhancement 5 Communication Optimization Among Multi-GPU 5.1 Challenges of Multi-GPU 5.2 Communication Scheme 5.3 Optimization of Communication 6 Performance Evaluation 6.1 Environment Setup 6.2 Performance on Single Node 6.3 Performance on Multiple Nodes 7 Conclusion References Approximation Algorithms for Reliability-Aware Maximum VoI on AUV-Aided Data Collections 1 Introduction 2 Related Works 2.1 AUV-Aided Data Collection 2.2 Orienteering Problem and Variants 3 System Model and Problem Definition 3.1 System Model 3.2 Problem Definition 4 Approximation Algorithm for the Path Finding Problem 4.1 Approximation Algorithm for the Path Finding Problem Without Real-Time VoI Decay 4.2 Approximation Algorithm for the Path Finding Problem with Real-Time VoI Decay 5 Simulation and Performance Evaluation 6 Conclusion References CCSBD: A Cost Control System Based on Blockchain and DRG Mechanism 1 Introduction 2 Related Work 3 System Design 3.1 System Overview 3.2 Medical Evidence-Based Classification Model 3.3 Contract Strategy for Clinical Data Sharing 4 Evaluation 5 Discussion 6 Conclusion References Number of UAVs and Mission Completion Time Minimization in Multi-UAV-Enabled IoT Networks 1 Introduction 2 Related Work 3 System Model and Problem Formulation 3.1 System Model 3.2 Problem Formulation 4 Algorithm Design 4.1 Optimization of UAV Number and Mission Allocation 4.2 Joint Optimization of Flight Time and Data Collection Time of UAV 5 Simulation Results 6 Conclusion References A Spatial-Temporal Similarity-Based Cooperative Surveillance Framework by Edge 1 Introduction 2 The Problem in Application 2.1 Spatial Similarity 2.2 Time Similarity 2.3 Spatiotemporal Similarity 3 Diversity Maximization Based on Spatiotemporal Similarity 4 Experiment 5 Conclusion References A Progressive Transmission Method of Cloud Point Data for HD Map in Autonomous Driving 1 Introduction 2 Progressive Transmission of Point Cloud Data 2.1 Data Compression for Progressive Transmission 2.2 Transfer Data and Restore Process 2.3 Progressive Transmission Time Model 3 Experiment 3.1 Experimental Set up 3.2 Comparison of Transmission Experiment Results 4 Conclusion References IMRSim: A Disk Simulator for Interlaced Magnetic Recording Technology 1 Introduction 2 IMR Simulation 2.1 IMRSim Kernel Module 2.2 User Interface 3 Simulation Analysis 3.1 Experimental Methodology 3.2 Experimental Results 4 Related Works 5 Conclusion References Storage and I/O Alleviating Performance Interference Through Intra-Queue I/O Isolation for NVMe-over-Fabrics 1 Introduction 2 Background and Motivation 2.1 NVMe-over-TCP 2.2 Intra-Queue Performance Interference 2.3 Motivation 3 System Overview 4 PINoF Design 4.1 Intra-queue I/O Isolation 4.2 Specific I/O Paths 4.3 Isolated Scheduling 5 Performance Evaluation 5.1 Experimental Setup 5.2 Evaluation Results 6 Conclusion References WALOR: Workload-Driven Adaptive Layout Optimization of Raft Groups for Heterogeneous Distributed Key-Value Stores 1 Introduction 2 Background and Motivation 2.1 Background 2.2 Motivation 3 WALOR 4 Evaluation 4.1 Experimental Setup 4.2 Overall Results 4.3 Impacts of Different Heterogeneous Configurations 4.4 Impacts of System Scale 5 Related Work 6 Conclusion References Efficient Data Placement for Zoned Namespaces (ZNS) SSDs 1 Introduction 2 Relate Work 3 ZNS SSD-Aware Data Placement 3.1 Lifetime-Based Data Insertion 3.2 Lifetime Variance-Aware Garbage Collection 4 Performance Evaluation 4.1 Setting 4.2 Results 5 Conclusion References SchedP: I/O-aware Job Scheduling in Large-Scale Production HPC Systems 1 Introduction 2 Related Work 3 Method Design 3.1 Per-Job I/O Pattern Prediction Mechanism 3.2 I/O-aware Pre-schedule Mechanism 4 Implementations 4.1 Training of the Per-Job I/O Pattern Prediction Model 4.2 I/O-aware Pre-schedule Mechanism 4.3 Integration into Slurm 5 Evaluations 5.1 Experiment Method 5.2 Results and Analysis 6 Conclusions References SpacKV: A Pmem-Aware Key-Value Separation Store Based on LSM-Tree 1 Introduction 2 Background 2.1 LSM-tree Based KV Store 2.2 Intel Optane Persistent Memory 3 Motivation 3.1 Write Amplification 3.2 Challenges of KV Separation on Persistent Memory 4 SpacKV Design 4.1 System Overview 4.2 Cascading Compaction Operations 4.3 Compaction-triggered Garbage Collection 4.4 PM-aware Optimizations 4.5 Recovery 5 Experiments and Evaluation 5.1 Experimental Setup 5.2 Overall Performance Evaluation 5.3 YCSB Evaluation 6 Related Work 7 Conclusion References Consistent and Efficient Batch Operations for NoSQL Databases with Hybrid Timestamp 1 Introduction 2 Background and Motivation 2.1 NoSQL Database and Transactional Support 2.2 Benefits and Cost of Batch Operations 3 Hybrid Timestamp Design 3.1 Timestamp Design and Architecture 3.2 Basic Algorithm 3.3 Server-Side Optimistic Execution 4 Evaluation 4.1 Exerimantal Setup 4.2 Operation Latency 4.3 Scalability 5 Related Work 6 Conclusion References Author Index
Similar books
Robotic Computing on FPGAs (Synthesis Lectures on Distributed Computing Theory)
2021 · PDF
Creating Autonomous Vehicle Systems
2017 · PDF
Engineering Autonomous Vehicles and Robots: The Dragonfly Modular-Based Approach (Wiley - IEEE)
2020 · PDF
MySQL® Notes for Professionals book
2018 · PDF
MrExcel 2022: Boosting Excel
2022 · PDF
MrExcel 2022: Boosting Excel
2022 · PDF
Session C11: Ancient Cultural Landscapes in South Europe – their Ecological Setting and Evolution, Session C22: Gardeners from South America, Session S04: Agro-Pastoralism and Early Metallurgy Sessions, Session WS29: The Idea of Enclosure in Recent Iberian Prehistory, Session C88: Rhytmes et causalites des dynamiques de l'anthropisation en Europe entre 6500 ET 500 BC: Hypotheses socio-culturelles et/ou climatiques: Proceedings of the XV UISPP World Congress (Lisbon 4-9 September 2006) / Actes du XV Congrès Mondial (Lisbonne 4-9 Septembre 2006) Vol.36
2010 · PDF
THE BRITISH ARMY IN INDIA: ITS PRESERVATION BY AN APPROPRIATE CLOTHING, HOUSING, LOCATING, RECREATIVE EMPLOYMENT, AND HOPEFUL ENCOURAGEMENT OF THE TROOPS. with AN APPENDIX ON INDIA : THE CLIMATE OP ITS HILLS ; THE DEVELOPMENT OF ITS RESODRCBS, INDUSTRY, AND ARTS ; THE ADMINISTRATION OF JUSTICE ; THE BLACK ACT ; THE PROGRESS OF CHRISTIANITY ; THE TRAFFIC IN OPIUM ; THE VALUE OF INDIA ; PERMANENT CAUSES OF DISAFFECTION, AND OF THE RECENT REBELLION ; THE TRADITIONARY POLICY; MISGOVERNMENT BY NATIVE RULERS ; ANNEXATIONS OF THEIR TERRITORY, ETC.
1858 · PDF