ENGLISH

Text Analysis with R: For Students of Literature

Book information

Publisher
Springer
Year
2020
ISBN
3030396428, 9783030396428
Language
english
Format
PDF
Filesize
3 MB (3348194 bytes)
Series
Quantitative Methods in the Humanities and Social Sciences
Edition
2
Pages
276\283
Time added
2020-03-31 06:00:22

Description

Now in its second edition, Text Analysis with Rprovides a practical introduction to computational text analysis using the open source programming language R. R is an extremely popular programming language, used throughout the sciences; due to its accessibility, R is now used increasingly in other research areas. In this volume, readers immediately begin working with text, and each chapter examines a new technique or process, allowing readers to obtain a broad exposure to core R procedures and a fundamental understanding of the possibilities of computational text analysis at both the micro and the macro scale. Each chapter builds on its predecessor as readers move from small scale “microanalysis” of single texts to large scale “macroanalysis” of text corpora, and each concludes with a set of practice exercises that reinforce and expand upon the chapter lessons. The book’s focus is on making the technical palatable and making the technical useful and immediately gratifying. Text Analysis with R is written with students and scholars of literature in mind but will be applicable to other humanists and social scientists wishing to extend their methodological toolkit to include quantitative and computational approaches to the study of text. Computation provides access to information in text that readers simply cannot gather using traditional qualitative methods of close reading and human synthesis. This new edition features two new chapters: one that introduces dplyr and tidyr in the context of parsing and analyzing dramatic texts to extract speaker and receiver data, and one on sentiment analysis using the syuzhet package. It is also filled with updated material in every chapter to integrate new developments in the field, current practices in R style, and the use of more efficient algorithms. Preface to the Second Edition Preface from the First Edition (Still Relevant) Contents About the Authors List of Figures List of Tables Part I Microanalysis 1 R Basics 1.1 Introduction 1.2 Download and Install R 1.3 Download and Install RStudio 1.4 Download the Supporting Materials 1.5 RStudio 1.6 Let's Get Started 1.7 Saving Commands and R Scripts 1.8 Assignment Operators 1.9 Practice References 2 First Foray into Text Analysis with R 2.1 Loading the First Text File 2.2 A Word About Warnings, Errors, Typos, and Crashes 2.3 Separate Content from Metadata 2.4 Reprocessing the Content 2.5 Beginning Some Analysis 2.6 Practice 3 Accessing and Comparing Word Frequency Data 3.1 Introduction 3.2 Start Up Code 3.3 Accessing Word Data 3.4 Recycling 3.5 Practice 4 Token Distribution and Regular Expressions 4.1 Introduction 4.2 Start Up Code 4.3 A Word About Coding Style 4.4 Dispersion Plots 4.5 Searching with grep 4.6 Practice Reference 5 Token Distribution Analysis 5.1 Cleaning the Workspace 5.2 Start Up Code 5.3 Identifying Chapter Breaks with grep 5.4 The for Loop and if Conditional 5.5 The for Loop in Eight Parts 5.5.1 5.5.2 5.5.3 5.5.4 5.5.5 5.5.6 5.5.7 5.5.8 5.6 Accessing and Processing List Items 5.6.1 rbind 5.6.2 More Recycling 5.6.3 apply 5.6.4 do.call (do dot call) 5.6.5 cbind 5.7 Practice 6 Correlation 6.1 Introduction 6.2 Start Up Code 6.3 Correlation Analysis 6.4 A Word About Data Frames 6.5 Testing Correlation with Randomization 6.6 Practice 7 Measures of Lexical Variety 7.1 Lexical Variety and the Type-Token Ratio 7.2 Start Up Code 7.3 Mean Word Frequency 7.4 Extracting Word Usage Means 7.5 Ranking the Values 7.6 Calculating the TTR inside lapply 7.7 A Further Use of Correlation 7.8 Practice Reference 8 Hapax Richness 8.1 Introduction 8.2 Start Up Code 8.3 sapply 8.4 An Inline Conditional Function 8.5 Practice 9 Do It KWIC 9.1 Introduction 9.2 Custom Functions 9.3 A Tokenization Function 9.4 Finding Keywords and Their Contextual Neighbors 9.5 Practice Reference 10 Do It KWIC(er) (and Better) 10.1 Getting Organized 10.2 Separating Functions for Reuse 10.3 User Interaction 10.4 readline 10.5 Building a Better KWIC Function 10.6 Fixing Some Problems 10.7 Practice Part II Metadata 11 Introduction to dplyr 11.1 Start Up Code 11.2 Using stack to Create a Data Frame 11.3 Installing and Loading dplyr 11.4 Using mutate, filter, arrange, and select 11.4.1 Mutate 11.4.2 filter 11.4.3 select 11.4.4 arrange 11.5 Practice 12 Parsing TEI XML 12.1 Introduction 12.2 The Text Encoding Initiative (TEI) 12.3 Parsing XML with R Using the Xml2 Package 12.4 Accessing the Textual Content 12.5 Calculating the Word Frequencies 12.6 Practice Reference 13 Parsing and Analyzing Hamlet 13.1 Background 13.2 Collecting the Speakers 13.3 Collecting the Speeches 13.4 A Better Pairing 13.5 Practice Reference 14 Sentiment Analysis 14.1 A Brief Overview 14.2 Loading syuzhet 14.3 Loading a Text 14.4 Getting Sentiment Values 14.5 Accessing Sentiment 14.6 Plotting 14.7 Smoothing 14.8 Computing Plot Similarity 14.9 Practice References Part III Macroanalysis 15 Clustering 15.1 Introduction 15.2 Corpus Ingestion 15.3 Custom Functions 15.4 Unsupervised Clustering and the Euclidean Metric 15.5 Converting an R List into a Data Matrix 15.6 Reshaping from Long to Wide Format 15.7 Preparing Data for Clustering 15.8 Clustering the Data 15.9 Practice Reference 16 Classification 16.1 Introduction 16.2 A Small Authorship Experiment 16.3 Text Segmentation 16.4 Reshaping from Long to Wide Format 16.5 Mapping the Data to the Metadata 16.6 Reducing the Feature Set 16.7 Performing the Classification with SVM 16.8 Practice Reference 17 Topic Modeling 17.1 Introduction 17.2 R and Topic Modeling 17.3 Text Segmentation and Preparation 17.4 The R Mallet Package 17.5 Simple Topic Modeling with a Standard Stop List 17.6 Unpacking the Model 17.7 Topic Visualization 17.8 Topic Coherence and Topic Probability 17.9 Practice References 18 Part of Speech Tagging and Named Entity Recognition 18.1 Pre-processing Text with a Part-of-Speech Tagger 18.2 Saving and Loading .Rdata Files 18.3 Topic Modeling the Noun Data 18.4 Named Entity Recognition 18.5 Practice Appendix A: Variable Scope Example Appendix B: The LDA Buffet Appendix C: Practice Exercise Solutions C.1 Solutions for Chap.1 C.2 Solutions for Chap.2 C.3 Solutions for Chap.3 C.4 Solutions for Chap.4 C.5 Solutions for Chap.5 C.6 Solutions for Chap.6 C.7 Solutions for Chap.7 C.8 Solutions for Chap.8 C.9 Solutions for Chap.9 C.10 Solutions for Chap.10 C.11 Solutions for Chap.11 C.12 Solutions for Chap.12 C.13 Solutions for Chap.13 C.14 Solutions for Chap.14 C.15 Solutions for Chap.15 C.16 Solutions for Chap.16 C.17 Solutions for Chap.17 C.18 Solutions for Chap.18 Index

Similar books