ENGLISH

Computational Linguistics and Intelligent Text Processing: 20th International Conference, CICLing 2019 La Rochelle, France, April 7–13, 2019 Revised Selected Papers, Part I

Book information

Publisher
Springer
Year
2023
ISBN
3031243366, 9783031243363
Language
english
Format
PDF
Filesize
27 MB (27884560 bytes)
Series
Lecture Notes in Computer Science, 13451
Pages
680\681
Time added
2023-02-27 12:24:15

Description

The two-volume set LNCS 13451 and 13452 constitutes revised selected papers from the CICLing 2019 conference which took place in La Rochelle, France, April 2019. The total of 95 papers presented in the two volumes was carefully reviewed and selected from 335 submissions. The book also contains 3 invited papers. The papers are organized in the following topical sections: General, Information extraction, Information retrieval, Language modeling, Lexical resources, Machine translation, Morphology, sintax, parsing, Name entity recognition, Semantics and text similarity, Sentiment analysis, Speech processing, Text categorization, Text generation, and Text mining. Preface Organization Contents – Part I Contents – Part II General Visual Aids to the Rescue: Predicting Creativity in Multimodal Artwork 1 Introduction 2 Related Work 3 Creativity in Advertising 4 Multimodal Creativity Dataset 5 Creativity Detection in Visual Modality 5.1 Bag of Visual Words 5.2 Observable Visual Features 6 Creativity Detecting Features in Linguistic Modality 7 Creativity Prediction Experiment 7.1 Multimodal Creativity Score 7.2 Unimodal Experiments 7.3 Multimodal Fusion Results 8 Conclusion References Knowledge-Based Techniques for Document Fraud Detection: A Comprehensive Study 1 Introduction 2 Receipt Dataset for Fraud Detection 3 Knowledge-Based Techniques for Document Fraud Detection 4 Knowledge-Based Fact-Checking Methods 4.1 Matrix Factorization Models 4.2 Geometric Models 4.3 Deep Learning Models 5 Experimental Setup 5.1 Evaluation 5.2 Data Pre-processing 5.3 Data Enrichment: External Verification 6 Results 7 Discussion and Conclusions References Exploiting Metonymy from Available Knowledge Resources 1 Introduction 2 Analysing Metonymy in WordNet and SUMO 3 Modelling Metonymy 3.1 Exploring the Mapping 3.2 Exploring the Ontology 4 Evaluation Framework and Results 5 Discussion 6 Conclusion and Future Work References Robust Evaluation of Language–Brain Encoding Experiments 1 Introduction 2 Human-Centered Evaluation of Computational Models 3 Datasets 3.1 Isolated Stimuli 3.2 Continuous Stimuli 4 Encoding Model 4.1 Language Model 4.2 Voxel Selection 5 Evaluation Experiments 5.1 Pairwise Evaluation 5.2 Voxel-Wise Evaluation 5.3 Representational Similarity Analysis 6 Discussion 7 Conclusions References Connectives with Both Arguments External: A Survey on Czech 1 Introduction 1.1 Part of Speech and The Position of Connectives 1.2 Connectives with Both Arguments External 2 Method, Data and Tools 3 Analysis 3.1 Wide and Narrow Connective ``scopes'' 3.2 Testing Criteria 3.3 Corpus Evidence 4 Conclusion References Recognizing Weak Signals in News Corpora 1 Introduction 2 Related Work 3 Materials and Methods 3.1 Dataset 3.2 Machine Learning 3.3 Unsupervised Learning 4 Conclusions References Low-Rank Approximation of Matrices for PMI-Based Word Embeddings 1 Introduction 2 Low-Rank Approximations of the PMI-matrix 3 Experimental Setup 3.1 Corpus 3.2 Training 3.3 Evaluation 4 Results 4.1 Similarity Task 4.2 Analogy Task 5 Discussion 6 Conclusion References Text Preprocessing for Shrinkage Regression and Topic Modeling to Analyse EU Public Consultation Data 1 Introduction 2 Related Work 3 The Data 4 Methods to Prepare the Inputs 4.1 Cleaning, Filtering and Tokenization 4.2 Lemmatization and Stemming 4.3 Test Corpora Generation 5 LASSO Regression 5.1 Methodology 5.2 Result Comparison 6 LDA Topic Modeling 6.1 Process 6.2 Result Comparison 7 Discussion 8 Conclusion References Intelligibility of Highly Predictable Polish Target Words in Sentences Presented to Czech Readers 1 Introduction 2 Hypothesis 3 Method 3.1 Design of the Web-Based Cloze Translation Experiment 3.2 Stimuli 4 Results 4.1 Comparison: Target Words With vs. Without Context 4.2 Different Categories of Target Words 4.3 Analysis of Wrong Responses 4.4 Correlations and Model 5 Discussion and Conclusion References Information Extraction Multi-lingual Event Identification in Disaster Domain 1 Introduction 1.1 Problem Definition 2 Related Work 3 Methodology 3.1 Mono-lingual Word Embedding Representation 3.2 Multi-lingual Word Embedding Representation 3.3 Baseline Model for Event Identification 3.4 Proposed Multilingual Event Identification Model 4 Datasets and Experiments 4.1 Datasets 4.2 Experimental Setup 5 Results and Analysis 5.1 Error Analysis 6 Conclusion and Future Works References Detection and Analysis of Drug Non-compliance in Internet Fora Using Information Retrieval Approaches 1 Introduction 2 Methods 2.1 Reference and Test Data 2.2 Categorization with Information Retrieval 2.3 Supervised Categorization with Machine Learning 3 Results and Discussion 3.1 Supervised Categorization and Evaluation with Information Retrieval 3.2 Supervised Categorization with Machine Learning 3.3 Unsupervised Detection with Information Retrieval 3.4 Comparison of the Two Categorization Approaches 3.5 Limitations of the Current Work 4 Conclusions References Char-RNN and Active Learning for Hashtag Segmentation 1 Introduction 2 Neural Model for Hashtag Segmentation 2.1 Sequence Labeling Approach 3 Dataset 3.1 Russian Dataset 3.2 English Dataset 4 Active Learning 5 Experiments 5.1 Baseline 5.2 Neural Model 5.3 Active Learning 5.4 Visualization 6 Related Work 6.1 Active Learning in NLP 6.2 Training on Synthetic Data 7 Conclusions References Extracting Food-Drug Interactions from Scientific Literature: Relation Clustering to Address Lack of Data 1 Introduction 2 Related Work 3 Dataset 4 Grouping Types of Relation 4.1 Intuitive Grouping 4.2 Unsupervised Clustering 5 Experiments 5.1 Clustering 5.2 FDI Type Classification 6 Results and Discussion 7 Conclusion and Future Work References Contrastive Reasons Detection and Clustering from Online Polarized Debates 1 Introduction 2 Related Work 3 Methodology 3.1 Phrase Mining Phase 3.2 Topic-Viewpoint Modeling Phase 3.3 Grouping and Facet Labeling 3.4 Reasons Table Extraction 4 Experiments and Results 4.1 Datasets 4.2 Experiments Set up 4.3 Evaluating Argument Facets Detection 4.4 Evaluating Informativeness 4.5 Evaluating Relevance and Clustering 5 Conclusion References Visualizing and Analyzing Networks of Named Entities in Biographical Dictionaries for Digital Humanities Research 1 Introduction 2 Extracting Named Entities from Biographical Texts 3 Applications 4 Assessment and Evaluation 5 Conclusions References Unsupervised Keyphrase Extraction from Scientific Publications 1 Introduction 2 Related Work 2.1 Keyphrase Extraction 2.2 Multivariate Outlier Detection Methods 3 Our Approach 3.1 Learning Vector Representations 3.2 Filtering Non-keyphrase Words 3.3 Generating Candidate Keyphrases 3.4 Scoring Candidate Keyphrases 4 Empirical Study 4.1 Experimental Setup 4.2 Evaluation Based on the Proportion of Outlier Vectors 4.3 Evaluation Based on the Type of Outlier Detection Method 4.4 Comparison with Other Approaches 4.5 Qualitative Results 5 Conclusions and Future Work References Information Retrieval Retrieving the Evidence of a Free Text Annotation in a Scientific Article: A Data Free Approach 1 Introduction 2 Related Work 3 Methods 3.1 GeneRIFs 3.2 Publications 3.3 Algorithm 3.4 Evaluation 4 Results and Discussion 4.1 Preliminary Results 4.2 Evidence Retrieval in Abstracts 4.3 Evidence Retrieval in Fulltexts 4.4 Missed GeneRIFs 5 Conclusion References Salience-Induced Term-Driven Serendipitous Web Exploration 1 The Framework 1.1 Queries 2 Information Objects 2.1 URLs 2.2 Terms 2.3 Synsets 2.4 Concepts 3 Stability of Web Search 4 Salience and Serendipity Lattices 4.1 Discussion 5 Conclusion and Future Works References Language Modeling Two-Phased Dynamic Language Model: Improved LM for Automated Language Translation 1 Introduction 1.1 Motivation 1.2 Related Work 2 Domain Adoption in the Language Model 2.1 Preprocessing 2.2 Base Corpus 2.3 Phase-One 2.4 Phase-Two 2.5 Final Language Model 3 Experimental Evaluation 4 Conclusions and Future Works References Composing Word Vectors for Japanese Compound Words Using Dependency Relations 1 Introduction 2 Related Work 3 Composing Word Vectors of Japanese Compound Words Using Dependency Relation 4 Compound Words and Dependency Relations 5 Experiment 6 Results 6.1 Unit Number of a Hidden Layer 6.2 Effect of Dependency Relations 6.3 Cosine Similarities of Each Dependency Relation 7 Discussion 7.1 Performances of the Models and Classification Accuracy of SVM 7.2 Error Analysis 7.3 Fine-Tuning of Models 8 Conclusions References Microtext Normalization for Chatbots 1 Introduction 2 Related Work 2.1 Microtext Analysis 2.2 Dialogue System 3 Proposed Framework for Chatbot 3.1 Datasets 4 Results and Discussion 4.1 Dataset Collection and Annotation 4.2 Time Complexity 4.3 BLEU Score 5 Conclusion and Future Work References Building Personalized Language Models Through Language Model Interpolation 1 Introduction 2 Related Work 3 Data and Evaluation 3.1 Dataset 3.2 Evaluation 3.3 Evaluation Metrics 4 Experimental Results 4.1 Tuning User-Trained Language Models 4.2 Impact of Volume of Data on User-Level Language Models 4.3 Interpolation 4.4 Accuracy at 3 Given c Keystrokes Evaluation 5 Conclusions References dpUGC: Learn Differentially Private Representation for User Generated Contents (Best Paper Award, Third Place, Shared) 1 Introduction 1.1 Goal of the Paper 1.2 Previous Work 2 Preliminaries 2.1 Differential Privacy 2.2 Word Embedding 3 Methodologies: Differentially Private Word Embedding 3.1 Differentially Private (DP-) Embedding 3.2 Personalised DP-Embedding 4 Experimental Settings 4.1 Evaluation Criteria 4.2 Datasets 4.3 Experiment Design 5 Evaluation Results 5.1 Evaluation #1: Changes in Semantic Space 5.2 Evaluation #2: Regression Task 6 Conclusions References Multiplicative Models for Recurrent Language Modeling 1 Introduction 2 Recurrent Neural Networks 3 Multiplicative RNNs 4 Sharing Intermediate States 4.1 mLSTM 4.2 True mLSTM 4.3 GRU 4.4 True mGRU 4.5 mGRU with Shared Intermediate State 5 Experiments in Character-Level Language Modeling 5.1 Penn Treebank 5.2 Text8 6 Conclusion References Impact of Gender Debiased Word Embeddings in Language Modeling 1 Introduction 2 Background 2.1 Words Embeddings 2.2 Debiased Word Embeddings 2.3 Recurrent Language Model 3 Research Questions 4 Experimental Framework 4.1 Datasets 4.2 Parameters 5 Results 6 Conclusions and Further Work References Initial Explorations on Chaotic Behaviors of Recurrent Neural Networks 1 Introduction 2 Chaotic Nature of Simple Vanilla RNN 2.1 Vanilla RNN in 1D Case 2.2 Multidimensional Case for Vanilla RNN 3 Chaotic Behavior of RHN 3.1 RHN Chaoticity in 1D 3.2 RHN Chaoticity in 2D 4 Experiments 5 Conclusion and Future Work References Lexical Resources LingFN: A Framenet for the Linguistic Domain 1 General Background and Introduction 2 The Data 3 The General Architecture of LingFN 3.1 Frame Types 3.2 Frame-to-Frame Relations 4 Methodology 4.1 Framenet Development 4.2 Frame Identification and Construction 4.3 Frame Element Identification and Development 4.4 Current Status of LingFN 5 Applications of LingFN 6 Conclusions and Future Work References SART - Similarity, Analogies, and Relatedness for Tatar Language: New Benchmark Datasets for Word Embeddings Evaluation 1 Introduction 2 Related Works 3 Proposed Datasets 3.1 Similarity Dataset 3.2 Relatedness Dataset 3.3 Analogies Dataset 4 Experiments and Evaluation 4.1 Similarity and Relatedness Results 4.2 Analogies Results 4.3 Comparison with English 5 Conclusion References Cross-Lingual Transfer for Distantly Supervised and Low-Resources Indonesian NER 1 Introduction 2 Related Works 2.1 Deep Character Embedding 2.2 Bidirectional Language Models (BiLM) 2.3 Cross-Lingual Transfer via Multi-task Learning 3 Proposed Method 3.1 Supervised Cross-Lingual Transfer with ELMo 3.2 Unsupervised Cross-Lingual Transfer via ELMo Fine-Tuning 4 Dataset 4.1 Gold Named Entity Corpus 4.2 Noisy Named Entity Corpus 4.3 ID-POS Corpus 4.4 Unlabeled Corpus for Language Model 5 Experiments 6 Results and Analysis 6.1 English Dataset Results 6.2 Indonesian Dataset Results 6.3 Cross-Lingual Transfer Analysis 7 Conclusion References Phrase-Level Simplification for Non-native Speakers 1 Introduction 2 Phrase-Level Substitution Generation 2.1 Retrofitted POS-Aware Phrase Embeddings 2.2 Generating Candidates with Phrase Embeddings 3 Comparison-Based Substitution Selection 4 Ranking with Phrase-Level Features 5 BenchPS: A New Dataset 5.1 Collecting Complex Replaceable Phrases 5.2 Collecting Phrase Simplifications 5.3 Simplification Ranking 6 Experiments 6.1 Candidate Phrase Generation 6.2 Phrase Simplicity Ranking 6.3 Full Pipeline Evaluation 6.4 Error Analysis 7 Final Remarks References Automatic Creation of a Pharmaceutical Corpus Based on Open-Data 1 Introduction 2 General Concepts 3 Related Work 4 ABA Corpus Construction 5 Analysis and Results 6 Conclusions 7 Future Work References Fool's Errand: Looking at April Fools Hoaxes as Disinformation Through the Lens of Deception and Humour 1 Introduction 2 Background 2.1 Deception Detection 2.2 Fake News 2.3 Humour Recognition 2.4 Irony 2.5 Satire 3 Hoax Feature Set 4 Data Collection 4.1 April Fools Corpus 4.2 News Corpus 4.3 Limitations 5 Analysis 5.1 Classifying April Fools' 5.2 Classifying ``Fake News'' 5.3 Individual Feature Performances 6 Conclusion References Russian Language Datasets in the Digital Humanities Domain and Their Evaluation with Word Embeddings 1 Introduction 2 Related Work 3 Tasks and Datasets 3.1 Task Types 3.2 Dataset Translation 3.3 Word Embedding Models 3.4 Implementation 4 Evaluation 4.1 Evaluation Setup 4.2 Evaluation Results 5 Discussion 6 Conclusions References Towards the Automatic Processing of Language Registers: Semi-supervisedly Built Corpus and Classifier for French 1 Introduction 2 State of the Art and Positioning 3 Proposed Approach 4 Data 5 Classifier Training 6 Automatically Labeled Corpus 7 Conclusion A Supplementary Material References Machine Translation Evaluating Terminology Translation in MT 1 Introduction 2 Related Work 2.1 Terminology Annotation 2.2 Term Translation Evaluation Method 2.3 PB-SMT Versus NMT: Terminology Translation & Evaluation 3 MT Systems 3.1 PB-SMT System 3.2 NMT System 3.3 Data Used 3.4 PB-SMT Versus NMT 4 Creating Gold-Standard Evaluation Set 4.1 Annotation Suggestions from Bilingual Terminology 4.2 Variations of Term 4.3 Consistency in Annotation 4.4 Ambiguity in Terminology Translation 4.5 Measuring Performance of TermMarker 5 Evaluating Terminology Translation in MT 5.1 Automatic Evaluation: TermEval 5.2 Term Translation Accuracy with TermEval 5.3 Manual Evaluation Method 5.4 Measuring Term Translation Accuracy from Manual Classification Results 5.5 Overlapping Correct and Incorrect Term Translations 5.6 Validating TermEval 6 Discussion and Analysis 6.1 False Positives 6.2 False Negatives 7 Conclusion References Detecting Machine-Translated Paragraphs by Matching Similar Words 1 Introduction 2 Related Work 2.1 Parsing Tree 2.2 N-gram Model 2.3 Word Distribution 2.4 Word Similarity 3 Proposed Method 3.1 Overview 3.2 Detail 4 Evaluation 4.1 Dataset 4.2 Comparison 4.3 Individual Features 4.4 Other Languages 5 Conclusion References Improving Low-Resource NMT with Parser Generated Syntactic Phrases 1 Introduction 2 Related Works 3 Proposed Method 3.1 Phrase Extraction from Parse Tree 3.2 Phrase Based SMT for TargetSource 3.3 Synthetic Parallel Corpus Using Back-Translation 3.4 Copied Parallel Corpus 3.5 NMT Training with Synthetic and Copied Corpus 4 Datasets 5 Experimental Setup 6 Results and Analysis 7 Conclusion References How Much Does Tokenization Affect Neural Machine Translation? 1 Introduction 2 Neural Machine Translation 3 Tokenizers 4 Experimental Framework 4.1 Corpora 4.2 Systems 4.3 Evaluation Metrics 5 Results 6 Qualitative Analysis 7 Conclusions References Take Help from Elder Brother: Old to Modern English NMT with Phrase Pair Feedback 1 Introduction 2 Related Works 3 Proposed Method 3.1 Overview of NMT 3.2 Phrase Augmentation 4 Data Sets 5 Experimental Setup 6 Results and Analysis 7 Conclusion References Adaptation of Machine Translation Models with Back-Translated Data Using Transductive Data Selection Methods 1 Introduction 2 Related Work 2.1 Transductive Data Selection Algorithms 2.2 Using Approximated Target Side 3 Fine-Tuning Models with Synthetic Data 4 Experiments 4.1 Experimental Settings 4.2 Model Adaptation with Subsets of Data 5 Results 5.1 Model Adaptation with Synthetic Data 5.2 Batch and Online Processing 6 Conclusion and Future Work References Morphology, Syntax, Parsing Automatic Detection of Parallel Sentences from Comparable Biomedical Texts 1 Introduction 2 Method 2.1 Comparable Corpora 2.2 Reference Data 2.3 Automatic Detection and Alignment of Parallel Sentences 2.4 Experimental Design 2.5 Evaluation 3 Presentation and Discussion of Results 4 Conclusion and Future Work References MorphBen: A Neural Morphological Analyzer for Bengali Language 1 Introduction 2 Related Work 2.1 Recent Models for Morphological Analysis 3 Our Neural Model for Morphological Analysis 3.1 Sentence Level Encoder 3.2 Word Level Predictor 4 Dataset, Experiments and Results 4.1 Data 4.2 Implementation Details 4.3 Morphological Analysis 4.4 PoS tagger 5 Conclusions References CCG Supertagging Using Morphological and Dependency Syntax Information 1 Introduction 2 From Dependency Syntax to CCG Derivation Tree 3 Machine Learning and Supertagging 4 Neural Network Model 4.1 Input Features 4.2 Basic Bi-directional LSTM and CRF Models 4.3 Our Model for Feature Set Enrichment and CCG Supertagging 5 Evaluation 5.1 Dataset and Preprocessing 5.2 Training Procedure 5.3 Experimental Results 6 Conclusion References Representing Overlaps in Sequence Labeling Tasks with a Novel Tagging Scheme: Bigappy-Unicrossy 1 Introduction 2 Related Work 3 Corpus 4 MWE Identification 4.1 Challenges 4.2 Evaluation Metrics 5 Tagging Schemes 5.1 IOB1 Tagging Scheme 5.2 IOB2 Tagging Scheme 5.3 Gappy 1-Level Tagging Scheme 5.4 Bigappy-Unicrossy Tagging Scheme 6 Model and Experiments 7 Results 8 Conclusion References *Paris is Rain. or It is raining in Paris?: Detecting Overgeneralization of Be-verb in Learner English 1 Introduction 2 Be-verb Sentence and Overgeneralization of be-verb 3 Proposed Method 3.1 Detection Procedure 3.2 Subject Complement Equivalence Check 3.3 How to Determine Hyperparameters 4 Evaluation 5 Discussion 6 Conclusions References Speeding up Natural Language Parsing by Reusing Partial Results 1 Introduction 2 Reuse of Partial Results 3 Generating the Templates 4 Experiments 4.1 Data and Evaluation 4.2 Model 4.3 Results 5 Ongoing Work 6 Conclusion References Unmasking Bias in News 1 Introduction 2 Investigating Masking for Hyperpartisanship Detection 3 Experiments 3.1 Masking Content vs. Style in Hyperpartisan News 3.2 Experimental Setup 3.3 Results and Discussion 4 Conclusions References Author Index

Similar books