Computer Vision – ECCV 2022: 17th European Conference, Tel Aviv, Israel, October 23–27, 2022, Proceedings, Part VI
Book information
Description
The 39-volume set, comprising the LNCS books 13661 until 13699, constitutes the refereed proceedings of the 17th European Conference on Computer Vision, ECCV 2022, held in Tel Aviv, Israel, during October 23–27, 2022. The 1645 papers presented in these proceedings were carefully reviewed and selected from a total of 5804 submissions. The papers deal with topics such as computer vision; machine learning; deep neural networks; reinforcement learning; object recognition; image classification; image processing; object detection; semantic segmentation; human pose estimation; 3d reconstruction; stereo vision; computational photography; neural networks; image coding; image reconstruction; object recognition; motion estimation. Foreword Preface Organization Contents – Part VI UnrealEgo: A New Dataset for Robust Egocentric 3D Human Motion Capture 1 Introduction 2 Related Work 2.1 Datasets for Outside-in 3D Human Pose Estimation 2.2 Datasets for Egocentric 3D Human Pose Estimation 2.3 Methods for Egocentric 3D Human Pose Estimation 3 UnrealEgo Dataset 3.1 Setup 3.2 Egocentric Dataset 4 Egocentric 3D Human Pose Estimation 4.1 2D Module 4.2 3D Module 5 Experiments 5.1 Implementation Details 5.2 Comparisons 5.3 Results 5.4 Ablation Study 6 Conclusions References Skeleton-Parted Graph Scattering Networks for 3D Human Motion Prediction 1 Introduction 2 Related Works 2.1 Human Motion Prediction 2.2 Graph Representation Learning 3 Skeleton-Parted Graph Scattering Network 3.1 Problem Formulation 3.2 Model Architecture 4 Multi-part Graph Scattering Block 4.1 Single-part Adaptive Graph Scattering 4.2 Bipartite Cross-Part Fusion 4.3 Loss Function 5 Experiments 5.1 Datasets 5.2 Model and Experimental Settings 5.3 Comparison to State-of-the-Art Methods 5.4 Model Analysis 6 Conclusion References Rethinking Keypoint Representations: Modeling Keypoints and Poses as Objects for Multi-person Human Pose Estimation 1 Introduction 2 Related Work 3 KAPAO: Keypoints and Poses as Objects 3.1 Architectural Details 3.2 Loss Function 3.3 Inference 3.4 Limitations 4 Experiments 4.1 Microsoft COCO Keypoints 4.2 CrowdPose 4.3 Ablation Studies 5 Conclusion References VirtualPose: Learning Generalizable 3D Human Pose Models from Virtual Data 1 Introduction 2 Related Work 3 Generalization Study 3.1 Baselines and Datasets 3.2 Experimental Results 4 VirtualPose 4.1 Abstract Geometry Representation 4.2 Root Estimation Network 4.3 Pose Estimation Network 5 Experiments 5.1 Implementation Details 5.2 Comparison to the State-of-the-arts 5.3 Ablation Study 5.4 Qualitative Results 6 Conclusion 6.1 Future Work References Poseur: Direct Human Pose Regression with Transformers 1 Introduction 2 Related Work 3 Method 3.1 Poseur Architecture 3.2 Training Targets and Loss Functions 3.3 Inference 4 Experiments 4.1 Implementation Details 4.2 Ablation Study 4.3 Extensions: End-to-End Pose Estimation 4.4 Main Results 5 Conclusion References SimCC: A Simple Coordinate Classification Perspective for Human Pose Estimation 1 Introduction 2 Related Work 3 SimCC: Reformulating HPE from Classification Perspective 3.1 Comparisons to 2D Heatmap-Based Approaches 4 Experiments 4.1 COCO Keypoint Detection 4.2 Ablation Study 4.3 CrowdPose 4.4 MPII Human Pose Estimation 5 Limitation and Future Work 6 Conclusion References Regularizing Vector Embedding in Bottom-Up Human Pose Estimation 1 Introduction 2 Related Work 2.1 Bottom-Up Methods 2.2 Vector Embedding 3 Our Method 3.1 Model Framework 3.2 Coupled Embedding 3.3 Improving Heatmap Regression with Coupled Embedding 3.4 Loss Function 4 Experiments 4.1 Datasets and Implementation Details 4.2 Comparison with SOTA 4.3 Group Margin 4.4 Comparison of Keypoint Grouping 4.5 Ablation Study 4.6 Hyper-parameter Study 5 Conclusions References A Visual Navigation Perspective for Category-Level Object Pose Estimation 1 Introduction 2 Related Work 3 Object Pose Estimation as Visual Navigation 3.1 Problem Formulation 3.2 Gradient Descent 3.3 Reinforcement Learning 3.4 Imitation Learning 4 Implementation Details 4.1 Pose-Aware Generative Model 4.2 Navigation Policy 5 Experiments 5.1 How Are Policies Affected by Design Choices? 5.2 What Is a Good Navigation Policy? 5.3 Comparison to the State-of-the-Art 6 Conclusions References Faster VoxelPose: Real-time 3D Human Pose Estimation by Orthographic Projection 1 Introduction 2 Related Work 2.1 Multi-view 3D Pose Estimation 2.2 Efficient Human Pose Estimation 3 Method 3.1 Overview 3.2 Human Detection Networks 3.3 Joint Localization Networks 4 Experiments 4.1 Setup 4.2 Evaluation and Comparison 4.3 Ablation Study 5 Conclusion References Learning to Fit Morphable Models 1 Introduction 2 Related Work 3 Method 3.1 Neural Fitter 3.2 Human Body Model and Fitting Tasks 3.3 Human Face Model and Fitting Task 3.4 Data Terms 3.5 Training Details 4 Experiments 4.1 Metrics 4.2 Quantitative Evaluation 4.3 Discussion 5 Conclusion References EgoBody: Human Body Shape and Motion of Interacting People from Head-Mounted Devices 1 Introduction 2 Related Work 3 Building the EgoBody Dataset 3.1 Interaction Scenarios 3.2 Data Acquisition Setup 3.3 Ground-truth Acquisition 4 EgoBody Dataset 5 Experiments 5.1 Benchmark Evaluation Metrics 5.2 Baseline Evaluation 5.3 Baseline Improvement 5.4 Cross-dataset Evaluation on You2Me 6 Conclusion References Grasp'D: Differentiable Contact-Rich Grasp Synthesis for Multi-Fingered Hands 1 Introduction 2 Related Work 3 Grasp'D: Differentiable Contact-rich Grasp Synthesis 3.1 Rigid Body Dynamics 3.2 Object Model with Coarse-to-fine Surface Smoothing 3.3 Contact Dynamics with Leaky Gradient 3.4 Grasping Metric and Problem Relaxation 3.5 Optimization 4 Experiments 4.1 Experimental Setup 4.2 Grasp Synthesis with ShapeNet Models 4.3 Grasp Synthesis from RGB-D Input of Unknown Objects 4.4 Ablation Study 5 Conclusions References AutoAvatar: Autoregressive Neural Fields for Dynamic Avatar Modeling 1 Introduction 2 Related Work 3 Method 3.1 Shape Encoding via Articulated Observer Points 3.2 Dynamic Feature Encoding 3.3 Articulation-Aware Shape Decoding 3.4 Implementation Details 4 Experimental Results 4.1 Datasets and Metrics 4.2 Evaluation 5 Conclusion References Deep Radial Embedding for Visual Sequence Learning 1 Introduction 2 Related Work 2.1 Connectionist Temporal Classification 2.2 Deep Feature Learning 3 Method 3.1 A Toy Sequence Recognition Example 3.2 The Design of Loss Function for Sequence Recognition 3.3 RadialCTC 4 Experiments 4.1 Datasets 4.2 Experimental Results 5 Conclusion References SAGA: Stochastic Whole-Body Grasping with Contact 1 Introduction 2 Related Work 3 Method 3.1 Overview 3.2 Whole-Body Grasping Pose Generation 3.3 Generative Motion Infilling 3.4 Contact-Aware Grasping Motion Optimization 4 Experiments 4.1 Stochastic Whole-Body Grasp Pose Synthesis 4.2 Stochastic Motion Infilling 4.3 Whole-Body Grasp Motion Synthesis 5 Conclusion and Discussion References Neural Capture of Animatable 3D Human from Monocular Video 1 Introduction 2 Related Works 3 Method 3.1 Mesh-Guided NeRF 3.2 Query Embedding for NeRF 3.3 Joint Mesh Estimation and NeRF Training 4 Experiments 4.1 Experimental Setup 4.2 Ablation Studies 4.3 Comparsions 4.4 Applications 5 Conclusion References General Object Pose Transformation Network from Unpaired Data 1 Introduction 2 Related Work 3 Method 4 Experiment 4.1 Ablation Studies 4.2 Qualitative Comparison 4.3 Multi-object Pose Transformation 4.4 Cross-Category Pose Transformation 4.5 Applications 4.6 Failure Case 5 More-In-Depth Discussion 6 Conclusion References Compositional Human-Scene Interaction Synthesis with Semantic Control 1 Introduction 2 Related Work 3 Method 3.1 Preliminaries 3.2 Interaction Synthesis with Semantic Control 3.3 Compositional Interaction Generation 4 Experiments 4.1 Dataset 4.2 Evaluation Metrics 4.3 Baselines 4.4 Interaction Synthesis with Semantic Control 4.5 Compositional Interaction Generation 5 Conclusion References PressureVision: Estimating Hand Pressure from a Single RGB Image*-6pt 1 Introduction 2 Related Work 3 The PressureVisionDB Dataset 3.1 The Capture Setup 4 Estimating Pressure and Contact from RGB 4.1 Network Architecture 4.2 Evaluation Metrics 4.3 Dataset Splits 5 Results 5.1 Baseline Models 5.2 Can Inference Succeed with New People? 5.3 Is Hand-Related Appearance Used for Inference? 5.4 Is Inference Reasonable with New Conditions? 6 Conclusion References PoseScript: 3D Human Poses from Natural Language 1 Introduction 2 Related Work 3 The PoseScript Dataset 3.1 Dataset Collection 3.2 Automatic Captioning Pipeline 3.3 Dataset Statistics 4 Application to Text-to-Pose Retrieval 5 Application to Text-Conditioned Pose Generation 6 Conclusion References DProST: Dynamic Projective Spatial Transformer Network for 6D Pose Estimation 1 Introduction 2 Related Work 3 Method 3.1 Framework Overview 3.2 Object Space Grid Generator 3.3 3D Feature Reconstruction 3.4 Projector and Pose Estimator 3.5 Objective Functions 4 Experiments 4.1 Datasets and Evaluation Metrics 4.2 Experiment Setup 4.3 Comparison with the State-of-the-Art 4.4 Ablation Studies 4.5 Runtime Analysis 5 Conclusion References 3D Interacting Hand Pose Estimation by Hand De-occlusion and Removal 1 Introduction 2 Related Work 2.1 Amodal Instance Segmentation and De-occlusion 2.2 Monocular RGB-Based Hand Pose Estimation 3 Method 3.1 Overview 3.2 Hand Amodal Segmentation Module (HASM) 3.3 Hand De-occlusion and Removal Module (HDRM) 3.4 3D Single Hand Pose Estimation (SHPE) 4 Amodal InterHand (AIH) Dataset 5 Experiments 5.1 Implementation Details 5.2 Datasets and Evaluation Metrics 5.3 Comparisons with State-of-the-Art Methods 5.4 Effect of Hand De-occlusion and Removal (HDR) Framework 5.5 Ablation Study 5.6 Time Complexity Analysis 5.7 Qualitative Results 6 Conclusions and Limitations References Pose for Everything: Towards Category-Agnostic Pose Estimation 1 Introduction 2 Related Works 2.1 2D Pose Estimation 2.2 Category-Agnostic Estimation 2.3 Few-Shot Learning 3 Class-Agnostic Pose Estimation (CAPE) 3.1 Problem Definition 3.2 POse Matching Network (POMNet) 4 Mulit-category Pose (MP-100) Dataset 5 Experiments 5.1 Implementation Details 5.2 Benchmark Results on MP-100 Dataset 5.3 Cross Super-Category Pose Estimation 5.4 Ablation Study 5.5 Qualitative Results 6 Conclusions and Limitations References PoseGPT: Quantization-Based 3D Human Motion Generation and Forecasting 1 Introduction 2 Related Work 3 The PoseGPT Model 3.1 Learning a Discrete Latent Space Representation 3.2 Learning a Density Model in the Discrete Latent Space 4 Experiments 4.1 Evaluation Metrics 4.2 Ablative Study of Design Choices 4.3 Comparison to the State of the Art 5 Conclusion References DH-AUG: DH Forward Kinematics Model Driven Augmentation for 3D Human Pose Estimation 1 Introduction 2 Related Work 3 Method 3.1 Overview 3.2 DH Parameter Model 3.3 Architecture 4 Experiments 4.1 Implementation Details 4.2 Datasets 4.3 Pose Augmentation in Video Pose Estimation 4.4 Pose Augmentation in Single-Frame Pose Estimation 4.5 Qualitative Results 4.6 Ablation Study 4.7 Limitation Analysis 5 Conclusion References Estimating Spatially-Varying Lighting in Urban Scenes with Disentangled Representation 1 Introduction 2 Related Work 3 Problem Formulation 4 Method 4.1 Dataset 4.2 Network Architecture 4.3 Training 4.4 Inference 5 Experiments 5.1 Ablation Study 5.2 Experimental Results 6 Conclusion References Boosting Event Stream Super-Resolution with a Recurrent Neural Network 1 Introduction 2 Related Work 3 Method 3.1 Problem Definition 3.2 Overall Pipeline 3.3 Recurrent Neural Network for Event SR 4 Experimental Results 5 Ablation Study 6 Downstream Event-Driven Applications 7 Conclusion References Projective Parallel Single-Pixel Imaging to Overcome Global Illumination in 3D Structure Light Scanning 1 Introduction 1.1 Contributions 2 Related Work 2.1 3D Reconstruction Under Global Illumination 2.2 Light Transport Coefficients Capture 3 Background 4 Projective Parallel Single-Pixel Imaging for Efficient Separation of Direct and Global Illumination 4.1 Local Maximum Constraint Proposition 4.2 Projective Single-Pixel Imaging for Projection Functions Capture 4.3 Local Slice Extension Method for Efficient Projection Functions Capture 5 Experiments and Evaluations 5.1 Compound Scene 5.2 Inter-reflections 5.3 Subsurface Scattering 5.4 Step Edges 6 Conclusion References Semantic-Sparse Colorization Network for Deep Exemplar-Based Colorization 1 Introduction 2 Related Work 3 Methods 3.1 Overview of the Proposed Method 3.2 Global Color Transfer 3.3 Local Details Transfer 3.4 Discussion 3.5 Objective Functions 4 Experiments 4.1 Implementation Details 4.2 Comparison with Previous Methods 5 Ablation Studies 6 Conclusions References Practical and Scalable Desktop-Based High-Quality Facial Capture 1 Introduction 2 Related Work 2.1 Active Illumination 2.2 Passive Capture 3 Desktop-Based Capture System 3.1 Tablet-Based Setup 3.2 Monitor-Based Setup 3.3 Modulated Binary Illumination 4 Reflectance and Shape Estimation 4.1 Acquisition Using White Illumination 4.2 Color-Multiplexed Illumination 4.3 Specular Roughness Estimation 4.4 Dynamic Capture 4.5 Base Geometry Acquisition 5 Results 5.1 Evaluation 5.2 Static Capture 5.3 Dynamic Capture 5.4 Limitations 6 Conclusions References FAST-VQA: Efficient End-to-End Video Quality Assessment with Fragment Sampling 1 Introduction 2 Related Works 3 Approach 3.1 Grid Mini-patch Sampling (GMS) 3.2 Fragment Attention Network (FANet) 4 Experiments 4.1 Evaluation Setup 4.2 Benchmark Results 4.3 Efficiency of FAST-VQA 4.4 Transfer Learning with Video-Quality-Related Representations 4.5 Ablation Studies on fragments 4.6 Ablation Studies on FANet 4.7 Reliability and Robustness Analyses 4.8 Qualitative Results: Local Quality Maps 5 Conclusions References Physically-Based Editing of Indoor Scene Lighting from a Single Image 1 Introduction 2 Related Work 3 Material and Light Source Prediction 3.1 Light Source Representation 3.2 Light Source Prediction 4 Neural Rendering Framework 4.1 Direct Shading Rendering Module 4.2 Depth-Based Hybrid Shadow Rendering Module 4.3 Indirect Shading Prediction 4.4 Predicting Lighting from Shading 4.5 Implementation Details 5 Experiments 6 Conclusions References LEDNet: Joint Low-Light Enhancement and Deblurring in the Dark 1 Introduction 2 Related Work 3 LOL-Blur Dataset 3.1 Existing Synthesis Methods and Limitations 3.2 Data Generation Pipeline 4 LEDNet 4.1 Low-Light Enhancement Encoder 4.2 Deblurring Decoder 4.3 Filter Adaptive Skip Connection 4.4 Loss Function 5 Experiments 5.1 Evaluation on LOL-Blur Dataset 5.2 Evaluation on Real Data 5.3 Ablation Study 6 Conclusion References MPIB: An MPI-Based Bokeh Rendering Framework for Realistic Partial Occlusion Effects 1 Introduction 2 Related Work 3 MPIB: An MPI-Based Bokeh Rendering Framework 3.1 MPI Representation and Layer Compositing Formulation 3.2 Background Inpainting Module 3.3 High-Resolution MPI Representation Module 3.4 Model Training 4 Experiments 4.1 Bokeh Rendering on Synthesized Dataset 4.2 Bokeh Rendering on Real-World Images 4.3 Ablation Study 5 Conclusion References Real-RawVSR: Real-World Raw Video Super-Resolution with a Benchmark Dataset 1 Introduction 2 Related Work 2.1 Image and Video SR Datasets 2.2 Image and Video SR Methods 3 Real-RawVSR Dataset Construction 4 The Proposed Method 4.1 Packing and Feature Extraction 4.2 Co-alignment 4.3 Interaction 4.4 Temporal Fusion 4.5 Channel Fusion 4.6 Reconstruction and Upsampling 4.7 Color Correction and Loss Function 5 Experiments 5.1 Training Details 5.2 Comparison with State-of-the-arts 5.3 Ablation Study 6 Conclusion and Discussion References Transform Your Smartphone into a DSLR Camera: Learning the ISP in the Wild 1 Introduction 2 Related Work 3 Method 3.1 ISP Network 3.2 Color Prediction 3.3 Color Mapping Module 3.4 Learning the Camera ISP 4 Dataset 5 Experiments 5.1 Ablative Analysis of the Color Mapping 5.2 Ablative Study of the Training Loss 5.3 Ablative Study of the Color Prediction Network 5.4 State-of-the-Art Comparison 6 Conclusion References Learning Deep Non-blind Image Deconvolution Without Ground Truths 1 Introduction 1.1 Problem Setting and Main Idea 1.2 Main Contributions 2 Related Works 3 Proposed Approach 3.1 Self-supervised Reconstruction Loss 3.2 Approximate Supervision in Image Space with Kernel Diversity and Cross-Image Patch Recurrence 3.3 Self-supervised Prediction Loss 3.4 Unsupervised Training and Ensemble Inference 3.5 NN Architecture 4 Performance Evaluation 4.1 Motion Deblurring with Erroneous Kernels 4.2 Motion Deblurring with Accurate Kernels 4.3 Microscopic Deconvolution 4.4 Ablation Studies 5 Conclusion References NEST: Neural Event Stack for Event-Based Image Enhancement 1 Introduction 2 Related Work 2.1 Event Representation 2.2 Event-Based Image Enhancement 3 NEST: Representation 3.1 Bidirectional Event Summation 3.2 Neural Representation 3.3 NEST Estimator 4 NEST: Application 4.1 NEST-Guided Image Deblurring 4.2 NEST-Guided Image Super-Resolution 4.3 NEST-Guided HFR Video Generation 4.4 Implementation Details 4.5 Ablation Study 5 Conclusion References Editable Indoor Lighting Estimation 1 Introduction 2 Related Work 3 Editable Indoor Lighting Representation 3.1 Lighting Representation 3.2 Ground Truth Dataset 3.3 Virtual Object Rendering 4 Approach 5 Experiments 5.1 Validation of Our 1-Light Approximation 5.2 Light Estimation Comparison 5.3 Ablation Study on Input Layout 5.4 Ablation Study on the Texture Network 6 Editing the Estimated Lighting 7 Discussion References Fast Two-Step Blind Optical Aberration Correction 1 Introduction 2 Related Work 3 Local PSF Parametric Model 3.1 Optical Aberrations Model 3.2 Blur Parametric Approximation 4 Proposed Method 4.1 Blind Gaussian Deblurring 4.2 Red and Blue Edge Correction 5 Experiments 5.1 Blind Grayscale PSF Removal 5.2 Lateral Chromatic Aberration Compensation 5.3 Real-World Examples 6 Conclusion References Seeing Far in the Dark with Patterned Flash 1 Introduction 2 Related Work 2.1 Flash Imaging 2.2 Low-Light Imaging Without Flash 2.3 Structured Light (SL) 3D Imaging 3 Image Formation Model 3.1 Signal-to-Noise Ratio Analysis 4 Patterned Flash Processing 4.1 Network Architecture 4.2 Loss Functions 5 Implementations 6 Simulation Results 6.1 Joint Image and Depth Estimation 6.2 Illumination Pattern Design 7 Experimental Results 8 Discussions and Conclusions References PseudoClick: Interactive Image Segmentation with Click Imitation 1 Introduction 2 Related Work 3 Method 3.1 Segmentation Error Decoder 3.2 Pseudo Clicks Generation 3.3 Pseudo Clicks Encoding 3.4 Loss Function 3.5 Implementation Details 4 Experiments 4.1 Evaluation Details 4.2 Comparison with State-of-the-Art 4.3 Cross-Domain Evaluation 4.4 Comparison Study 5 Limitations 6 Conclusion References Author Index
Similar books
Computer Vision – ECCV 2022: 17th European Conference, Tel Aviv, Israel, October 23–27, 2022, Proceedings, Part VIII
2022 · PDF
Computer Vision – ECCV 2022: 17th European Conference, Tel Aviv, Israel, October 23–27, 2022, Proceedings, Part IX (Lecture Notes in Computer Science)
Advanced Topics in Computer Vision (Advances in Computer Vision and Pattern Recognition)
2014 · EPUB
Computer Vision, Imaging and Computer Graphics Theory and Applications. 16th International Joint Conference, VISIGRAPP 2021 Virtual Event, February 8–10, 2021 Revised Selected Papers
2023 · PDF
Computer Vision, Imaging and Computer Graphics Theory and Applications: 16th International Joint Conference, VISIGRAPP 2021 Virtual Event, February 8–10, 2021 Revised Selected Papers
2023 · PDF
Pattern Recognition: ICPR International Workshops and Challenges: Virtual Event, January 10–15, 2021, Proceedings, Part IV
2021 · PDF
Dense Image Correspondences for Computer Vision
2015 · PDF
Machine Learning, Optimization, and Big Data: First International Workshop, MOD 2015, Taormina, Sicily, Italy, July 21-23, 2015, Revised Selected Papers
2015 · PDF