Text Mining with Machine Learning: Principles and Techniques
Book information
Description
This book provides a perspective on the application of machine learning-based methods in knowledge discovery from natural languages texts. By analysing various data sets, conclusions which are not normally evident, emerge and can be used for various purposes and applications. The book provides explanations of principles of time-proven machine learning algorithms applied in text mining together with step-by-step demonstrations of how to reveal the semantic contents in real-world datasets using the popular R-language with its implemented machine learning algorithms. The book is not only aimed at IT specialists, but is meant for a wider audience that needs to process big sets of text documents and has basic knowledge of the subject, e.g. e-mail service providers, online shoppers, librarians, etc. The book starts with an introduction to text-based natural language data processing and its goals and problems. It focuses on machine learning, presenting various algorithms with their use and possibilities, and reviews the positives and negatives. Beginning with the initial data pre-processing, a reader can follow the steps provided in the R-language including the subsuming of various available plug-ins into the resulting software tool. A big advantage is that R also contains many libraries implementing machine learning algorithms, so a reader can concentrate on the principal target without the need to implement the details of the algorithms her- or himself. To make sense of the results, the book also provides explanations of the algorithms, which supports the final evaluation and interpretation of the results. The examples are demonstrated using realworld data from commonly accessible Internet sources. Cover Title Page Copyright Page Dedication Preface Contents Authors’ Biographies 1. Introduction to Text Mining with Machine Learning 1.1 Introduction 1.2 Relation of Text Mining to Data Mining 1.3 The Text Mining Process 1.4 Machine Learning for Text Mining 1.4.1 Inductive Machine Learning 1.5 Three Fundamental Learning Directions 1.5.1 Supervised Machine Learning 1.5.2 Unsupervised Machine Learning 1.5.3 Semi-supervised Machine Learning 1.6 Big Data 1.7 About This Book 2. Introduction to R 2.1 Installing R 2.2 Running R 2.3 RStudio 2.3.1 Projects 2.3.2 Getting Help 2.4 Writing and Executing Commands 2.5 Variables and Data Types 2.6 Objects in R 2.6.1 Assignment 2.6.2 Logical Values 2.6.3 Numbers 2.6.4 Character Strings 2.6.5 Special Values 2.7 Functions 2.8 Operators 2.9 Vectors 2.9.1 Creating Vectors 2.9.2 Naming Vector Elements 2.9.3 Operations with Vectors 2.9.4 Accessing Vector Elements 2.10 Matrices and Arrays 2.11 Lists 2.12 Factors 2.13 Data Frames 2.14 Functions Useful in Machine Learning 2.15 Flow Control Structures 2.15.1 Conditional Statement 2.15.2 Loops 2.16 Packages 2.16.1 Installing Packages 2.16.2 Loading Packages 2.17 Graphics 3. Structured Text Representations 3.1 Introduction 3.2 The Bag-of-Words Model 3.3 The Limitations of the Bag-of-Words Model 3.4 Document Features 3.5 Standardization 3.6 Texts in Different Encodings 3.7 Language Identification 3.8 Tokenization 3.9 Sentence Detection 3.10 Filtering Stop Words, Common, and Rare Terms 3.11 Removing Diacritics 3.12 Normalization 3.12.1 Case Folding 3.12.2 Stemming and Lemmatization 3.12.3 Spelling Correction 3.13 Annotation 3.13.1 Part of Speech Tagging 3.13.2 Parsing 3.14 Calculating the Weights in the Bag-of-Words Model 3.14.1 Local Weights 3.14.2 Global Weights 3.14.3 Normalization Factor 3.15 Common Formats for Storing Structured Data 3.15.1 Attribute-Relation File Format (ARFF) 3.15.2 Comma-Separated Values (CSV) 3.15.3 C5 format 3.15.4 Matrix Files for CLUTO 3.15.5 SVMlight Format 3.15.6 Reading Data in R 3.16 A Complex Example 4. Classification 4.1 Sample Data 4.2 Selected Algorithms 4.3 Classifier Quality Measurement 5. Bayes Classifier 5.1 Introduction 5.2 Bayes’ Theorem 5.3 Optimal Bayes Classifier 5.4 Naïve Bayes Classifier 5.5 Illustrative Example of Naïve Bayes 5.6 Naïve Bayes Classifier in R 5.6.1 Running Naïve Bayes Classifier in RStudio 5.6.2 Testing with an External Dataset 5.6.3 Testing with 10-Fold Cross-Validation 6. Nearest Neighbors 6.1 Introduction 6.2 Similarity as Distance 6.3 Illustrative Example of k-NN 6.4 k-NN in R 7. Decision Trees 7.1 Introduction 7.2 Entropy Minimization-Based c5 Algorithm 7.2.1 The Principle of Generating Trees 7.2.2 Pruning 7.3 C5 Tree Generator in R 7.3.1 Generating a Tree 7.3.2 Information Acquired from C5-Tree 7.3.3 Using Testing Samples to Assess Tree Accuracy 7.3.4 Using Cross-Validation to Assess Tree Accuracy 7.3.5 Generating Decision Rules 8. Random Forest 8.1 Introduction 8.1.1 Bootstrap 8.1.2 Stability and Robustness 8.1.3 Which Tree Algorithm? 8.2 Random Forest in R 9. Adaboost 9.1 Introduction 9.2 Boosting Principle 9.3 Adaboost Principle 9.4 Weak Learners 9.5 Adaboost in R 10. Support Vector Machines 10.1 Introduction 10.2 Support Vector Machines Principles 10.2.1 Finding Optimal Separation Hyperplane 10.2.2 Nonlinear Classification and Kernel Functions 10.2.3 Multiclass SVM Classification 10.2.4 SVM Summary 10.3 SVM in R 11. Deep Learning 11.1 Introduction 11.2 Artificial Neural Networks 11.3 Deep Learning in R 12. Clustering 12.1 Introduction to Clustering 12.2 Difficulties of Clustering 12.3 Similarity Measures 12.3.1 Cosine Similarity 12.3.2 Euclidean Distance 12.3.3 Manhattan Distance 12.3.4 Chebyshev Distance 12.3.5 Minkowski Distance 12.3.6 Jaccard Coefficient 12.4 Types of Clustering Algorithms 12.4.1 Partitional (Flat) Clustering 12.4.2 Hierarchical Clustering 12.4.3 Graph Based Clustering 12.5 Clustering Criterion Functions 12.5.1 Internal Criterion Functions 12.5.2 External Criterion Function 12.5.3 Hybrid Criterion Functions 12.5.4 Graph Based Criterion Functions 12.6 Deciding on the Number of Clusters 12.7 K-Means 12.8 K-Medoids 12.9 Criterion Function Optimization 12.10 Agglomerative Hierarchical Clustering 12.11 Scatter-Gather Algorithm 12.12 Divisive Hierarchical Clustering 12.13 Constrained Clustering 12.14 Evaluating Clustering Results 12.14.1 Metrics Based on Counting Pairs 12.14.2 Purity 12.14.3 Entropy 12.14.4 F-Measure 12.14.5 Normalized Mutual Information 12.14.6 Silhouette 12.14.7 Evaluation Based on Expert Opinion 12.15 Cluster Labeling 12.16 A Few Examples 13. Word Embeddings 13.1 Introduction 13.2 Determining the Context and Word Similarity 13.3 Context Windows 13.4 Computing Word Embeddings 13.5 Aggregation of Word Vectors 13.6 An Example 14. Feature Selection 14.1 Introduction 14.2 Feature Selection as State Space Search 14.3 Feature Selection Methods 14.3.1 Chi Squared (x(sup[x])) 14.3.2 Mutual Information 14.3.3 Information Gain 14.4 Term Elimination Based on Frequency 14.5 Term Strength 14.6 Term Contribution 14.7 Entropy-Based Ranking 14.8 Term Variance 14.9 An Example References Index
Similar books
MySQL® Notes for Professionals book
2018 · PDF
MrExcel 2022: Boosting Excel
2022 · PDF
MrExcel 2022: Boosting Excel
2022 · PDF
Session C11: Ancient Cultural Landscapes in South Europe – their Ecological Setting and Evolution, Session C22: Gardeners from South America, Session S04: Agro-Pastoralism and Early Metallurgy Sessions, Session WS29: The Idea of Enclosure in Recent Iberian Prehistory, Session C88: Rhytmes et causalites des dynamiques de l'anthropisation en Europe entre 6500 ET 500 BC: Hypotheses socio-culturelles et/ou climatiques: Proceedings of the XV UISPP World Congress (Lisbon 4-9 September 2006) / Actes du XV Congrès Mondial (Lisbonne 4-9 Septembre 2006) Vol.36
2010 · PDF
THE BRITISH ARMY IN INDIA: ITS PRESERVATION BY AN APPROPRIATE CLOTHING, HOUSING, LOCATING, RECREATIVE EMPLOYMENT, AND HOPEFUL ENCOURAGEMENT OF THE TROOPS. with AN APPENDIX ON INDIA : THE CLIMATE OP ITS HILLS ; THE DEVELOPMENT OF ITS RESODRCBS, INDUSTRY, AND ARTS ; THE ADMINISTRATION OF JUSTICE ; THE BLACK ACT ; THE PROGRESS OF CHRISTIANITY ; THE TRAFFIC IN OPIUM ; THE VALUE OF INDIA ; PERMANENT CAUSES OF DISAFFECTION, AND OF THE RECENT REBELLION ; THE TRADITIONARY POLICY; MISGOVERNMENT BY NATIVE RULERS ; ANNEXATIONS OF THEIR TERRITORY, ETC.
1858 · PDF
Idries Shah 27 Books Collection : A Perfumed Scorpion, A Veiled Gazelle, Caravan of Dreams, Darkest England, Destination Mecca, Evenings with Idries Shah, Knowing How to Know, Learning How to Learn, Letters and Lectures of Idries Shah, Neglected aspects of Sufi study, Observations, Oriental Magic, Reflections, Seeker after Truth, Special Illumination, Special Problems in the study of Sufi ideas, Sufi thought and action, Tales of the Dervishes, The Dermis Probe, The Elephant in the Dark, The Englishman Handbook, Idries Shah Antology, The Magic Monastery, The natives are restless, wisdom of the Idiots PDF.
2022 · PDF
The travels of Capts. Lewis and Clarke from St. Louis, by way of the Missouri and Columbia rivers, to the Pacific ocean; performed in the years 1804, 1805 & 1806, by order of the government of the United States. Containing delineations of the manners, customs, religion, &c. of the Indians, comp. from various authentic sources, and original documents, and a summary of the Statistical view of the Indian nations, from the official communication of Meriwether Lewis. Illustrated with a map of the country, inhabited by the western tribes of Indians
1809 · PDF
Professional Linux kernel architecture ''Wrox programmer to programmer''--Cover. - ''What you are reading right now is the result of an evolution over more than seven years: After two years of writing, the first edition was published in German by Carl Hanser Verlag in 2003. It then described kernel 2.6.0. The test was used as a basis for the low-level design documentation for the EAL4+ security evaluation of Red Hat Enterprise Linux 5, requiring to update it to kernel 2.6.18 (if the EAL acronym does not mean anything to you, then Wikipedia is once more your friend). Hewlett-Packard sponsored the translation into English and has, thankfully, granted the rights to publish the result. Updates to kernel 2.6.24 were then performed specifically for this book''--P. ix
2008 · PDF