Computer Vision – ECCV 2022: 17th European Conference, Tel Aviv, Israel, October 23–27, 2022, Proceedings, Part VIII
Book information
Description
The 39-volume set, comprising the LNCS books 13661 until 13699, constitutes the refereed proceedings of the 17th European Conference on Computer Vision, ECCV 2022, held in Tel Aviv, Israel, during October 23–27, 2022. The 1645 papers presented in these proceedings were carefully reviewed and selected from a total of 5804 submissions. The papers deal with topics such as computer vision; machine learning; deep neural networks; reinforcement learning; object recognition; image classification; image processing; object detection; semantic segmentation; human pose estimation; 3d reconstruction; stereo vision; computational photography; neural networks; image coding; image reconstruction; object recognition; motion estimation. Foreword Preface Organization Contents – Part VIII ECCV Caption: Correcting False Negatives by Collecting Machine-and-Human-verified Image-Caption Associations for MS-COCO 1 Introduction 2 Related Works 2.1 Noisy Many-to-Many Correspondences of Image-Caption Datasets 2.2 Machine-in-the-Loop (MITL) Annotation 3 ECCV Caption Dataset Construction 3.1 Model Candidates for Machine Annotators 3.2 Crowdsourcing on Amazon Mechanical Turk 3.3 Postprocessing MTurk Annotations 4 Re-evaluation of ITM Models on ECCV Caption 4.1 Evaluation Metrics and Comparison Methods 4.2 Re-evaluation of ITM Methods 5 Discussion and Limitations 6 Conclusion References MOTCOM: The Multi-Object Tracking Dataset Complexity Metric 1 Introduction 2 Related Work 3 Challenges in Multi-object Tracking 4 The MOTCOM Metrics 4.1 Occlusion Metric 4.2 Motion Metric 4.3 Visual Similarity Metric 4.4 MOTCOM 5 Evaluation 5.1 Ground Truth 5.2 Evaluation Metrics 6 Results 7 Discussion 8 Conclusion References How to Synthesize a Large-Scale and Trainable Micro-Expression Dataset? 1 Introduction 2 Related Work 3 Preliminaries 4 Synthesizing Micro-Expressions 4.1 The Proposed Protocol 4.2 Major Finding: Action Units that Constitute Trainable MiEs 4.3 The MiE-X Dataset 5 Experiment 5.1 Experimental Setups 5.2 Effectiveness of the Synthetic Database 5.3 Further Analysis 5.4 Understanding of MiEs: A Discussion 6 Conclusion References A Real World Dataset for Multi-view 3D Reconstruction 1 Introduction 2 Related Work 3 Data Acquisition 4 Data Annotation 4.1 Notations 4.2 Texture-rich Object Annotation 4.3 Textureless Object Annotation 5 Dataset Statistics 6 Evaluation 6.1 Experiments 7 Discussion 8 Conclusion References REALY: Rethinking the Evaluation of 3D Face Reconstruction 1 Introduction 2 Related Work 3 Background 3.1 Notation & Preliminaries 3.2 3DMM and Face Reconstruction 4 Motivation 5 REALY: A New 3D Face Benchmark 6 A Novel Evaluation Pipeline 7 Experiment 7.1 Ablation Study: bICPv.s. gICP 7.2 Evaluating Face Reconstruction Methods 7.3 Evaluating Different 3DMMs 8 Conclusions References Capturing, Reconstructing, and Simulating: The UrbanScene3D Dataset 1 Introduction 2 Related Work 3 The UrbanScene3D Dataset 4 Scene Acquisition with Aerial Path Planning 4.1 Aerial Path Planning Methods 4.2 Geometric Proxies 5 Scene Reconstruction Benchmarks 5.1 High-Precision LiDAR Scan 5.2 UAV Capturing Cost 5.3 Aerotriangulation Error 5.4 Reconstruction Accuracy and Completeness 5.5 Comparison of Different Planners 6 Simulator and Applications 7 Conclusion and Future Work References 3D CoMPaT: Composition of Materials on Parts of 3D Things 1 Introduction 2 Related Work 3 3D CoMPaT: Data Collection, Benchmark, and Validation 3.1 3D CAD Models Collection 3.2 Materials Collection 3.3 Part-Material Assignment 3.4 Rendering Composition of Materials on Parts of the Collected CAD Models 3.5 Dataset Statistics 3.6 Dataset Split and Non-compositional Validation Experiments 4 2D/3D Grounded CoMPaT Recognition (GCR) Task, Baselines, and Results 5 Conclusion References PartImageNet: A Large, High-Quality Dataset of Parts 1 Introduction 2 Related Work 2.1 Part-based Models 2.2 Part Datasets 3 PartImageNet Dataset 3.1 Data Collection 3.2 Annotation 3.3 Annotation Quality and Consistency 3.4 Statistics 4 Experiments 4.1 Semantic Part Segmentation 4.2 Object Segmentation 4.3 Few-shot Learning 5 Conclusion References A-OKVQA: A Benchmark for Visual Question Answering Using World Knowledge 1 Introduction 2 Related Work 3 A-OKVQA Collection 4 Dataset Statistics 5 Experiments 5.1 Evaluation 5.2 Large-scale Pre-trained Models 5.3 Rationale Generation 5.4 Specialized Models 6 Analysis of Models 7 Conclusion References OOD-CV: A Benchmark for Robustness to Out-of-Distribution Shifts of Individual Nuisances in Natural Images 1 Introduction 2 Related Works 3 Dataset Collection 3.1 What Are Important Nuisance Factors? 3.2 Collecting Images 3.3 Data Annotation 4 Experiments 4.1 Robustness to Individual Nuisances 4.2 Data Augmentation for Enhancing Robustness 4.3 Effect of Model Architecture on Robustness 4.4 OOD Shifts in Multiple Nuisances 5 Conclusion References Facial Depth and Normal Estimation Using Single Dual-Pixel Camera 1 Introduction 2 Related Work 3 Overview 4 Dual-Pixel Facial Dataset 4.1 Dataset Configuration 4.2 Hardware Setup 4.3 Ground Truth Data Acquisition 5 Facial Depth and Normal Estimation 5.1 Overall Architecture 5.2 Adaptive Sampling Module 5.3 Adaptive Normal Module 5.4 Depth and Normal Estimation 5.5 Implementation Details 6 Experiments 6.1 Comparison Results 6.2 Ablation Study 7 Conclusion References The Anatomy of Video Editing: A Dataset and Benchmark Suite for AI-Assisted Video Editing 1 Introduction 2 Related Works 3 Anatomy of Video Editing: Dataset 3.1 Shot Attributes 3.2 Scene Composition and Camera Setups 3.3 Annotation Procedure 3.4 Dataset Statistics 4 Anatomy of Video Editing: Benchmark Suite 4.1 Shot Attributes Classification 4.2 Camera Setup Clustering 4.3 Shot Sequence Ordering 4.4 Next Shot Selection 4.5 Missing Shot Attributes Prediction 5 Experimental Results and Discussion 5.1 Experimental Results 6 Conclusion References StyleBabel: Artistic Style Tagging and Captioning 1 Introduction 2 Related Work 3 StyleBabel Dataset 3.1 Study Context 3.2 StyleBabel Grounded Annotation 4 Visual Embedding (ALADIN-ViT) 5 StyleBabel Experimental Setup 6 Experiments and Discussion 6.1 Style Auto-tagging (style2text) 6.2 Style Description (style2text) 6.3 Text Based Style Retrieval (text2style) 7 Conclusion References PANDORA: A Panoramic Detection Dataset for Object with Orientation 1 Introduction 2 Related Work 2.1 Existing Bounding Boxes 2.2 Panoramic Object Detection Dataset 2.3 Panoramic Image Object Detection 3 RBFoV 3.1 RBFoV Representation 3.2 IoU Calculation Between Two RBFoVs 4 PANDORA Dataset 4.1 Image Collection 4.2 Category Selection 4.3 Image Annotation 4.4 Dataset Statistics 5 R-CenterNet 5.1 Network Architecture and Loss Definition 5.2 Implementation Details 5.3 Panoramic Rotation Data Augmentation 6 Experiment 6.1 Quantify BFoV and RBFoV 6.2 Evaluations 6.3 Ablation Study 7 Conclusion References FS-COCO: Towards Understanding of Freehand Sketches of Common Objects in Context 1 Introduction 2 Related Work 3 Dataset Collection 4 Dataset Composition 4.1 Comparison to Existing Datasets 5 Towards Scene Sketch Understanding 5.1 Semi-synthetic Versus Freehand Sketches 5.2 What Does a Freehand Sketch Capture? 5.3 Sketch Captioning 6 Efficient ``Pretext'' Task 6.1 Proposed Hierarchical Decoder (H-Decoder) 6.2 Evaluation and Discussion 7 Conclusion References Exploring Fine-Grained Audiovisual Categorization with the SSW60 Dataset 1 Introduction 2 Related Work 2.1 Image, Audio, and Video Datasets 2.2 Multi-modal Learning 3 SSW60 Dataset 4 Methods 4.1 Implementation Details 5 Experiments 6 Conclusion References The Caltech Fish Counting Dataset: A Benchmark for Multiple-Object Tracking and Counting 1 Introduction 2 Related Work 3 Dataset 4 Metrics 4.1 Counting Protocol 4.2 Counting Metric 4.3 Detection and Tracking Metrics 5 Experiments 5.1 Baseline 5.2 Ablation Study and Generalization Upper Bounds 5.3 Baseline++ 6 Conclusions References A Dataset for Interactive Vision-Language Navigation with Unknown Command Feasibility 1 Introduction 2 Related Work 3 MoTIF Dataset 3.1 Data Collection 3.2 Dataset Analysis 4 Task Feasibility Experiments 4.1 Models 4.2 Results 5 Task Automation Experiments 5.1 Models 5.2 Results 6 Discussion 7 Conclusion References BRACE: The Breakdancing Competition Dataset for Dance Motion Synthesis 1 Introduction 2 Related Work 2.1 Dance Datasets 2.2 Dance Motion Synthesis 3 The BRACE Dataset 3.1 Data Acquisition 3.2 Quality Control 3.3 Dance Elements 3.4 Audio Correlation 4 Testing Generative Models on BRACE 4.1 Evaluation Metrics 4.2 Results 5 Pose Estimation on BRACE 6 Conclusion References Dress Code: High-Resolution Multi-category Virtual Try-On*-6pt 1 Introduction 2 Related Work 3 Dress Code Dataset 4 Virtual Try-On with Pixel-Level Semantics 4.1 Baseline Architecture 4.2 Pixel-Level Semantic-Aware Discriminator 5 Experiments 5.1 Experimental Setup 5.2 Experiments on Dress Code 5.3 Experiments on VITON 6 Conclusion References A Data-Centric Approach for Improving Ambiguous Labels with Combined Semi-supervised Classification and Clustering 1 Introduction 1.1 Related Work 2 Method 2.1 Definitions 2.2 DC3 3 Experiments 3.1 Datasets 3.2 Metrics 3.3 Implementation Details 3.4 Evaluation 3.5 Proof-of-Concept Improved Data Quality 4 Discussion 5 Conclusion References ClearPose: Large-scale Transparent Object Dataset and Benchmark 1 Introduction 2 Related Works 2.1 Transparent Dataset and Annotation 2.2 Transparent Depth Completion and Object Pose Estimation 3 Dataset 3.1 Dataset Objects and Statistics 3.2 Pose Annotation 4 Benchmark Experiments 4.1 Depth Completion 4.2 Instance-Level Object Pose Estimation 5 Discussions 6 Conclusions References When Deep Classifiers Agree: Analyzing Correlations Between Learning Order and Image Statistics 1 Introduction 2 Problem Statement and Motivation 3 Agreement for Different Batch-Sizes and Architectures, as Well as for Random Labels 4 Dataset Metrics 5 Do Basic Image Statistics Correlate with Training Learning Dynamics? 6 Discussion on Limitations and Prospects 7 Conclusions and Outlook References AnimeCeleb: Large-Scale Animation CelebHeads Dataset for Head Reenactment 1 Introduction 2 Animation CelebHeads Dataset 2.1 Data Creation Process 2.2 Dataset Description 2.3 Animation Head Reenactment 3 Cross-domain Head Reenactment 3.1 Driving Pose Representations 3.2 Training Pipeline 3.3 Experiments 4 Conclusions References MUGEN: A Playground for Video-Audio-Text Multimodal Understanding and GENeration 1 Introduction 2 Related Work 3 MUGEN Dataset 4 Video-Audio-Text Retrieval and Generation 4.1 Video-Audio-Text Retrieval 4.2 Video-Audio-Text Generation 5 Experiments 5.1 Video-Audio-Text Retrieval 5.2 Video-Audio-Text Generation 6 Conclusion References A Dense Material Segmentation Dataset for Indoor and Outdoor Scene Parsing 1 Introduction 2 Related Work 3 Data Collection 3.1 Material Labels 3.2 Image Selection 3.3 Segmentation and Instances 3.4 Labeling 3.5 Label Fusion 4 Experiments 4.1 Cross-Dataset Comparison 4.2 Recognition of Different Skin Types 4.3 A Material Segmentation Benchmark 4.4 Real-World Examples 5 Discussion and Conclusion 6 Conclusion References MimicME: A Large Scale Diverse 4D Database for Facial Expression Analysis 1 Introduction 2 Related Work 2.1 3D/4D Face Datasets 2.2 3DMM and Blendshape Models 2.3 Appearance Models 3 MimicMe Database 3.1 Data Acquisition 3.2 Annotation and Registration Process 3.3 Texture Completion 3.4 Creating Expression Blendshapes 4 Experiments 4.1 Facial Expression Recognition 4.2 Evaluation of the Expression Blendshape Model 4.3 Non-linear 3D Expression Model 5 Conclusion References Delving into Universal Lesion Segmentation: Method, Dataset, and Benchmark 1 Introduction 2 Related Work 3 SegLesion Dataset 3.1 Data Collection and Annotation 3.2 Data Statistics 3.3 Potential Applications 3.4 Data Naming 4 Methodology 4.1 Knowledge Embedding Module 4.2 Knowledge Embedding Network 5 Experiments 5.1 Experimental Setup 5.2 Ablation Studies 5.3 Performance Comparison 6 Conclusion References Large Scale Real-World Multi-person Tracking 1 Introduction 2 Related Work 3 Video Sourcing 4 Annotation Pipeline 5 Dataset 5.1 Video-Level Statistics 5.2 Track-Level Statistics 5.3 Benchmarking 6 Experiments 6.1 Model Evaluation 6.2 In-Depth Model Analysis 7 Conclusion and Discussion References D2-TPred: Discontinuous Dependency for Trajectory Prediction Under Traffic Lights 1 Introduction 2 Related Work 3 D2-TPred 3.1 Problem Formulation 3.2 Spatio-Temporal Dependency 3.3 Trajectory Prediction Near Traffic Lights 3.4 VTP-TL Dataset 4 Experimental Evaluation 4.1 Quantitative Evaluation 4.2 Ablation Studies 4.3 Qualitative Evaluation 5 Conclusions References The Missing Link: Finding Label Relations Across Datasets 1 Introduction 2 Related Work 3 Method 3.1 Discovering Relations Using Visual Information 3.2 Relation Type Discovery 3.3 Predicting Relation Types Using Language 3.4 Discovering Relations by Combining Vision and Language 3.5 Evaluation 4 Results 5 Applications 5.1 Understand Label Relations 5.2 Identify Missing Aspects 5.3 Increase Label Specificity 6 Conclusion References Learning Omnidirectional Flow in 360 Video via Siamese Representation 1 Introduction 2 Related Work 3 FLOW360 Dataset 4 SLOF 5 Experiments 6 Conclusion References VizWiz-FewShot: Locating Objects in Images Taken by People with Visual Impairments 1 Introduction 2 Related Work 3 VizWiz-FewShot Dataset 3.1 Dataset Creation 3.2 Dataset Analysis 4 Algorithm Benchmarking 4.1 Few-Shot Instance Segmentation Algorithms 4.2 Few-Shot Object Detection 5 Conclusions References TRoVE: Transforming Road Scene Datasets into Photorealistic Virtual Environments 1 Introduction 2 Related Work 3 Our Approach 3.1 Building the Virtual Environment 3.2 Object and Camera Placement 3.3 Textures, Lighting and Background 3.4 Data Processing and Training 4 Experiments and Results 4.1 Datasets and Experiments 4.2 Result Analysis 4.3 Dataset Statistics 5 Conclusions References Trapped in Texture Bias? A Large Scale Comparison of Deep Instance Segmentation 1 Introduction 2 Methods 2.1 An Object-Centric Version of Stylized COCO 2.2 Model Selection 3 Related Work 4 Results 5 Discussion 6 Conclusion References Deformable Feature Aggregation for Dynamic Multi-modal 3D Object Detection 1 Introduction 2 Related Work 2.1 Object Detection with Point Cloud 2.2 Multi-modal 3D Object Detection 3 AutoAlignV2 3.1 Deformable Feature Aggregation 3.2 Depth-Aware GT-AUG 3.3 Image-Level Dropout Training Strategy 4 Experiments 4.1 Dataset and Experimental Setup 4.2 Main Results 4.3 Ablation Studies 4.4 Dynamic Inference and Runtime 5 Conclusion References WeLSA: Learning to Predict 6D Pose from Weakly Labeled Data Using Shape Alignment 1 Introduction 2 Related Work 3 Method 3.1 Architecture 3.2 Training Pipeline and Loss Functions 4 Results 4.1 Training Data 4.2 LineMOD Dataset 4.3 LineMOD-Occlusion Dataset 4.4 TLess Dataset 5 Ablation Studies 5.1 Amount of Training Data 5.2 Feature Decoder 5.3 Shape Network 5.4 Influence of Each Training Stage on the Final Performance 6 Conclusion References Graph R-CNN: Towards Accurate 3D Object Detection with Semantic-Decorated Local Graph 1 Introduction 2 Related Works 3 Methods 3.1 Dynamic Point Aggregation 3.2 RoI-Graph Pooling 3.3 Visual Features Augmentation 3.4 Loss Functions 4 Experiments 4.1 Datasets 4.2 Implementation Settings 4.3 Comparison with State-of-the-Art Methods 4.4 Ablation Study 5 Conclusions References MPPNet: Multi-frame Feature Intertwining with Proxy Points for 3D Temporal Object Detection*-6pt 1 Introduction 2 Related Work 3 Methodology 3.1 Single-Frame Proposal Network and 3D Proposal Trajectories 3.2 Three-Hierarchy Feature Aggregation with Proxy Points 3.3 Temporal 3D Detection Head and Optimization 4 Experiments 4.1 Dataset and Implementation Details 4.2 Main Results of MPPNet on Waymo Open Dataset 4.3 Ablation Studies 5 Conclusions References Long-tail Detection with Effective Class-Margins 1 Introduction 2 Related Works 3 Preliminary 4 Effective Class-Margins 4.1 Detection Error Bound 4.2 Ranking Bounds 4.3 Effective Class-Margin Loss 5 Experiments 5.1 Experimental Settings 5.2 Implementation Details 5.3 Experimental Results 6 Conclusion References Semi-supervised Monocular 3D Object Detection by Multi-view Consistency 1 Introduction 2 Related Work 2.1 Monocular 3D Object Detection 2.2 Semi-supervised Object Detection 2.3 Self-supervised Learning with Multi-view Data 3 Background 4 Approach 4.1 Box-Level Consistency 4.2 Object-Level Consistency 4.3 Overall Loss 5 Experiments 5.1 Datasets 5.2 Experimental Setup 5.3 Experimental Results on the KITTI Validation Set 5.4 Comparison with State-of-the-Art Detectors on the KITTI Test Set 5.5 Experimental Results on the nuScenes Dataset 5.6 Ablation Study 6 Conclusion References PTSEFormer: Progressive Temporal-Spatial Enhanced TransFormer Towards Video Object Detection 1 Introduction 2 Related Works 2.1 Vision Transformer 2.2 Video Object Detection 3 PTSEFormer 3.1 Overview 3.2 Temporal and Spatial Encoding 3.3 Enhanced Memory Decoding 3.4 Learning PTSEFormer 4 Experiments 4.1 Implement Details 4.2 State-of-the-Art Comparison 4.3 Ablation Studies 4.4 Visualization 5 Conclusion References Author Index
Similar books
Computer Vision – ECCV 2022: 17th European Conference, Tel Aviv, Israel, October 23–27, 2022, Proceedings, Part VI
2022 · PDF
Computer Vision – ECCV 2022: 17th European Conference, Tel Aviv, Israel, October 23–27, 2022, Proceedings, Part IX (Lecture Notes in Computer Science)
Advanced Topics in Computer Vision (Advances in Computer Vision and Pattern Recognition)
2014 · EPUB
Computer Vision, Imaging and Computer Graphics Theory and Applications. 16th International Joint Conference, VISIGRAPP 2021 Virtual Event, February 8–10, 2021 Revised Selected Papers
2023 · PDF
Computer Vision, Imaging and Computer Graphics Theory and Applications: 16th International Joint Conference, VISIGRAPP 2021 Virtual Event, February 8–10, 2021 Revised Selected Papers
2023 · PDF
Pattern Recognition: ICPR International Workshops and Challenges: Virtual Event, January 10–15, 2021, Proceedings, Part IV
2021 · PDF
Dense Image Correspondences for Computer Vision
2015 · PDF
Machine Learning, Optimization, and Big Data: First International Workshop, MOD 2015, Taormina, Sicily, Italy, July 21-23, 2015, Revised Selected Papers
2015 · PDF