Practical Text Analytics
Book information
Description
Preface......Page 3 Contents......Page 5 Abbreviations......Page 11 Figures......Page 12 Tables......Page 18 1.2 Text Analytics: What Is It?......Page 20 1.3 Origins and Timeline of Text Analytics......Page 22 1.4 Text Analytics in Business and Industry......Page 24 1.5 Text Analytics Skills......Page 25 1.7.1 Planning......Page 27 1.7.3 Text Analysis Techniques......Page 28 References......Page 29 Planning the Project......Page 31 2.1 Introduction......Page 32 2.2.2 Content Analysis for Inductive Inference......Page 33 2.3 Unitizing and the Unit of Analysis......Page 35 2.3.2 The Recording Unit......Page 36 2.5 Coding and Categorization......Page 37 2.6.1 Inductive Inference......Page 40 2.6.2 Deductive Inference......Page 41 Further Reading......Page 42 3.1 Introduction......Page 43 3.2 Initial Planning Considerations......Page 44 3.2.2 Objectives......Page 45 3.3 Planning Process......Page 46 3.4.1 Identifying the Analysis Problem......Page 47 3.5.1 Definition of the Project’s Scope and Purpose......Page 48 3.5.2 Text Data Collection......Page 49 3.5.3.1 Non-probability Sampling......Page 51 3.5.3.3 Sampling for Classification Analysis......Page 52 3.5.3.4 Sample Size......Page 53 3.6.1 Analysis Method Selection......Page 54 3.6.2 The Selection of Implementation Software......Page 55 References......Page 56 Further Reading......Page 57 Text Preparation......Page 58 4.1 Introduction......Page 59 4.3 Unitize and Tokenize......Page 60 4.3.1 N-Grams......Page 62 4.5 Stop Word Removal......Page 64 4.6.1 Syntax and Semantics......Page 67 4.6.2 Stemming......Page 68 4.6.3 Lemmatization......Page 70 4.6.4 Part-of-Speech (POS) Tagging......Page 72 Further Reading......Page 73 5.2 The Inverted Index......Page 74 5.3 The Term-Document Matrix......Page 77 5.4 Term-Document Matrix Frequency Weighting......Page 78 5.4.1 Local Weighting......Page 79 5.4.1.2 Binary/Boolean Frequency......Page 80 5.4.2.1 Document Frequency (df)......Page 81 5.4.2.2 Global Frequency (gf)......Page 82 5.4.2.3 Inverse Document Frequency (idf)......Page 83 5.4.3 Combinatorial Weighting: Local and Global Weighting......Page 84 5.4.3.1 Term Frequency-Inverse Document Frequency (tfidf)......Page 85 Further Reading......Page 86 Text Analysis Techniques......Page 87 6.1 Introduction......Page 88 6.2.1 Singular Value Decomposition (SVD)......Page 90 6.2.2 LSA Example......Page 91 6.3 Cosine Similarity......Page 95 6.4 Queries in LSA......Page 98 6.5 Decision-Making: Choosing the Number of Dimensions......Page 99 Further Reading......Page 102 7.1 Introduction......Page 103 7.2 Distance and Similarity......Page 104 7.3 Hierarchical Cluster Analysis......Page 108 7.3.2.1 Single Linkage......Page 109 7.3.2.2 Complete Linkage......Page 110 7.3.3.2 Ward’s Minimum Variance Method......Page 111 7.4 k-Means Clustering......Page 113 7.4.2 The kMC Process......Page 114 7.4.3 Advantages and Disadvantages of kMC......Page 118 7.5.1.1 Subjective Methods......Page 119 Silhouette Plot......Page 121 7.5.2 Naming/Describing Clusters......Page 122 7.5.3 Evaluating Model Fit......Page 123 7.5.4 Choosing the Cluster Analysis Model......Page 124 Further Reading......Page 125 8.1 Introduction......Page 126 8.2 Latent Dirichlet Allocation (LDA)......Page 128 8.3 Correlated Topic Model (CTM)......Page 129 8.4 Dynamic Topic Model (DT)......Page 131 8.5 Supervised Topic Model (sLDA)......Page 132 8.7.1 Assessing Model Fit and Number of Topics......Page 133 8.7.2 Model Validation and Topic Identification......Page 135 8.7.3 When to Use Topic Models......Page 137 References......Page 138 Further Reading......Page 139 9.1 Introduction......Page 140 9.3.1 Confusion Matrices/Contingency Tables......Page 141 9.3.2.2 Error Rate......Page 143 9.3.3.1 Precision......Page 144 9.3.3.3 F-Measure......Page 145 9.4.1 Naïve Bayes......Page 146 9.4.2 k-Nearest Neighbors (kNN)......Page 147 9.4.3 Support Vector Machines (SVM)......Page 149 9.4.4 Decision Trees......Page 150 9.4.5 Random Forests......Page 152 9.4.6 Neural Networks......Page 153 9.5.1 Model Fit......Page 155 References......Page 157 Further Reading......Page 158 10 Modeling Text Sentiment: Learning & Lexicon Models......Page 159 10.1 Lexicon Approach......Page 160 10.2 Machine Learning Approach......Page 166 10.2.1 Naïve Bayes (NB)......Page 167 10.2.2 Support Vector Machines (SVM)......Page 168 10.2.3 Logistic Regression......Page 169 10.3 Sentiment Analysis Performance: Considerations and Evaluation......Page 170 Further Reading......Page 172 Communicating the Results......Page 173 11.1 Introduction......Page 174 11.2 Telling Stories About the Data......Page 175 11.3.1 Storytelling Framework......Page 177 11.3.2 Applying the Framework......Page 178 11.4.2 Zillow......Page 180 Further Reading......Page 182 12 Visualizing Analysis Results......Page 183 12.1.3 Solidify the Message......Page 184 12.1.5 Keep It Simple......Page 185 12.2.1 Corpus/Document Collection-Level Visualizations......Page 186 12.2.2.2 Cluster-Level Visualizations......Page 188 12.2.2.3 Topic-Level Visualizations......Page 189 12.2.2.4 Category or Class-Level Visualizations......Page 191 12.2.2.5 Sentiment-Level Visualizations......Page 192 12.2.3 Document-Level Visualizations......Page 194 Further Reading......Page 196 Text Analytics Examples......Page 197 13.1 Introduction to R and RStudio......Page 198 13.2 SA Data and Data Import......Page 199 13.3 Objective of the Sentiment Analysis......Page 202 13.4.1 Tokenize......Page 204 13.4.2 Remove Stop Words......Page 205 13.5 Sentiment Analysis......Page 206 13.6 Sentiment Analysis Results......Page 209 13.7 Custom Dictionary......Page 211 13.8 Out-of-Sample Comparison......Page 222 References......Page 224 Further Reading......Page 225 14.1 Introduction to Python and IDLE......Page 226 14.2 Preliminary Steps......Page 227 14.3 Getting Started......Page 229 14.4 Data and Data Import......Page 230 14.5 Analysis......Page 232 Further Reading......Page 247 15.1 Introduction......Page 248 15.2 Getting Started in RapidMiner......Page 249 15.3 Text Data Import......Page 252 15.4 Text Preparation and Preprocessing......Page 254 15.5 Text Classification Sentiment Analysis......Page 261 Further Reading......Page 266 16.1 Introduction......Page 267 16.2 Getting Started......Page 268 16.3 Analysis......Page 270 Further Reading......Page 286 Index......Page 287
Similar books
Practical Text Analytics: Maximizing the Value of Text Data
2019 · PDF
Perilous Policing: Criminal Justice in Marginalized Communities
2019 · PDF
The Internet of People, Things and Services: Workplace Transformations
2018 · PDF
e-Research Collaboration: Theory, Techniques and Challenges
2010 · PDF
Aligning Business Strategies and Analytics: Bridging Between Theory and Practice
2019 · PDF
e-Research Collaboration: Theory, Techniques and Challenges
2010 · PDF
Personal Web Usage in the Workplace: A Guide to Effective Human Resources Management
2003 · CHM
MySQL® Notes for Professionals book
2018 · PDF