The Big R-Book. From Data Science to Learning Machines and Big Data
Book information
Description
Introduces professionals and scientists to statistics and machine learning using the programming language R Written by and for practitioners, this book provides an overall introduction to R, focusing on tools and methods commonly used in data science, and placing emphasis on practice and business use. It covers a wide range of topics in a single volume, including big data, databases, statistical machine learning, data wrangling, data visualization, and the reporting of results. The topics covered are all important for someone with a science/math background that is looking to quickly learn several practical technologies to enter or transition to the growing field of data science. The Big R-Book for Professionals: From Data Science to Learning Machines and Reporting with R includes nine parts, starting with an introduction to the subject and followed by an overview of R and elements of statistics. The third part revolves around data, while the fourth focuses on data wrangling. Part 5 teaches readers about exploring data. In Part 6 we learn to build models, Part 7 introduces the reader to the reality in companies, Part 8 covers reports and interactive applications and finally Part 9 introduces the reader to big data and performance computing. It also includes some helpful appendices. Provides a practical guide for non-experts with a focus on business users Contains a unique combination of topics including an introduction to R, machine learning, mathematical models, data wrangling, and reporting Uses a practical tone and integrates multiple topics in a coherent framework Demystifies the hype around machine learning and AI by enabling readers to understand the provided models and program them in R Shows readers how to visualize results in static and interactive reports Supplementary materials includes PDF slides based on the book’s content, as well as all the extracted R-code and is available to everyone on a Wiley Book Companion Site The Big R-Book is an excellent guide for science technology, engineering, or mathematics students who wish to make a successful transition from the academic world to the professional. It will also appeal to all young data scientists, quantitative analysts, and analytics professionals, as well as those who make mathematical models. Cover Short Overview Contents Foreword Preface About the Companion Site I Introduction 1 The Big Picture with Kondratiev and Kardashev 2 The ScientificMethod and Data 3 Conventions II Starting with R and Elements of Statistics 4 The Basics of R 4.1 GettingStartedwithR 4.2 Variables 4.3 DataTypes 4.3.1 The Elementary Types 4.3.2 Vectors 4.3.2.1 CreatingVectors 4.3.3 Accessing Data from a Vector 4.3.3.1 Vector Arithmetic 4.3.3.2 Vector Recycling 4.3.3.3 Reordering and Sorting 4.3.4 Matrices 4.3.4.1 Creating Matrices 4.3.4.2 Naming Rows and Columns 4.3.4.3 Access Subsets of a Matrix 4.3.4.4 Matrix Arithmetic 4.3.5 Arrays 4.3.5.1 Creating and Accessing Arrays 4.3.5.2 Naming Elements of Arrays 4.3.5.3 Manipulating Arrays 4.3.5.4 Applying Functions over Arrays 4.3.6 Lists 4.3.6.1 Creating Lists 4.3.6.2 Naming Elements of Lists 4.3.6.3 List Manipulations 4.3.7 Factors 4.3.7.1 Creating Factors 4.3.7.2 Ordering Factors 4.3.8 DataFrames 4.3.8.1 Introduction to Data Frames 4.3.8.2 Accessing Information from a Data Frame 4.3.8.3 Editing Data in a Data Frame 4.3.8.4 Modifying Data Frames 4.3.9 Strings or the Character-type 4.4 Operators 4.4.1 Arithmetic Operators 4.4.2 Relational Operators 4.4.3 Logical Operators 4.4.4 Assignment Operators 4.4.5 Other Operators 4.5 Flow Control Statements 4.5.1 Choices 4.5.1.1 The if-Statement 4.5.1.2 The Vectorised If-statement 4.5.1.3 The Switch-statement 4.5.2 Loops 4.5.2.1 The For Loop 4.5.2.2 Repeat 4.5.2.3 While 4.5.2.4 Loop Control Statements 4.6 Functions 4.6.1 Built-inFunctions 4.6.2 Help with Functions 4.6.3 User-defined Functions 4.6.4 Changing Functions 4.6.5 Creating Function with Default Arguments 4.7 Packages 4.7.1 Discovering Packages in R 4.7.2 Managing Packages in R 4.8 Selected Data Interfaces 4.8.1 CSV Files 4.8.2 Excel Files 4.8.3 Databases 5 Lexical Scoping and Environments 5.1 Environments in R 5.2 Lexical Scoping in R 6 The Implementation of OO 6.1 Base Types 6.2 S3 Objects 6.2.1 Creating S3 Objects 6.2.2 Creating Generic Methods 6.2.3 Method Dispatch 6.2.4 Group Generic Functions 6.3 S4 Objects 6.3.1 Creating S4 Objects 6.3.2 Using S4 Objects 6.3.3 Validation of Input 6.3.4 Constructor functions 6.3.5 The.Data slot 6.3.6 Recognising Objects, Generic Functions, and Methods 6.3.7 Creating S4 Generics 6.3.8 Method Dispatch 6.4 The Reference Class, refclass, RC or R5 Model 6.4.1 Creating RC Objects 6.4.2 Important Methods and Attributes 6.5 Conclusions about the OO Implementation 7 Tidy R with the Tidyverse 7.1 The Philosophy of the Tidyverse 7.2 Packages in the Tidyverse 7.2.1 The Core Tidyverse 7.2.2 The Non-core Tidyverse 7.3 Working with the Tidyverse 7.3.1 Tibbles 7.3.2 Piping with R 7.3.3 Attention Points When Using the Pipe 7.3.4 Advanced Piping 7.3.4.1 The Dollar Pipe 7.3.4.2 The T-Pipe 7.3.4.3 The Assignment Pipe 7.3.5 Conclusion 8 Elements of Descriptive Statistics 8.1 Measures of Central Tendency 8.1.1 Mean 8.1.1.1 The Arithmetic Mean 8.1.1.2 Generalised Means 8.1.2 The Median 8.1.3 The Mode 8.2 Measures of Variation or Spread 8.3 Measures of Covariation 8.3.1 The Pearson Correlation 8.3.2 The Spearman Correlation 8.3.3 Chi-square Tests 8.4 Distributions 8.4.1 Normal Distribution 8.4.2 Binomial Distribution 8.5 Creating an Overview of Data Characteristics 9 Visualisation Methods 9.1 Scatterplots 9.2 Line Graphs 9.3 Pie Charts 9.4 Bar Charts 9.5 Boxplots 9.6 Violin Plots 9.7 Histograms 9.8 Plotting Functions 9.9 Maps and Contour Plots 9.10 Heat-maps 9.11 Text Mining 9.11.1 Word Clouds 9.11.2 WordAssociations 9.12 Colours in R 10 Time Series Analysis 10.1 Time Series in R 10.1.1 The Basics of Time Series in R 10.1.1.1 The Function ts() 10.1.1.2 Multiple Time Series in one Object 10.2 Forecasting 10.2.1 Moving Average 10.2.1.1 The Moving Average in R 10.2.1.2 Testing the Accuracy of the Forecasts 10.2.1.3 Basic Exponential Smoothing 10.2.1.4 Holt-Winters Exponential Smoothing 10.2.2 Seasonal Decomposition 11 Further Reading III Data Import 12 A Short History ofModern Database Systems 13 RDBMS 14 SQL 14.1 Designing the Database 14.2 Building the Database Structure 14.2.1 Installing a RDBMS 14.2.2 Creating the Database 14.2.3 Creating the Tables and Relations 14.3 Adding Data to the Database 14.4 Querying the Database 14.4.1 The Basic Select Query 14.4.2 More Complex Queries 14.5 Modifying the Database Structure 14.6 Selected Features of SQL 14.6.1 Changing Data 14.6.2 Functions in SQL 15 Connecting R to an SQL Database IV Data Wrangling 16 Anonymous Data 17 DataWrangling in the tidyverse 17.1 Importing the Data 17.1.1 Importing from an SQL RDBMS 17.1.2 Importing Flat Files in the Tidyverse 17.1.2.1 CSV Files 17.1.2.2 Making Sense of Fixed-width Files 17.2 Tidy Data 17.3 Tidying Up Data with tidyr 17.3.1 Splitting Tables 17.3.2 Convert Headers to Data 17.3.3 Spreading One Column Over Many 17.3.4 Split One Columns into Many 17.3.5 Merge Multiple Columns Into One 17.3.6 Wrong Data 17.4 SQL-like Functionality via dplyr 17.4.1 Selecting Columns 17.4.2 Filtering Rows 17.4.3 Joining 17.4.4 Mutating Data 17.4.5 Set Operations 17.5 String Manipulation in the tidyverse 17.5.1 Basic String Manipulation 17.5.2 Pattern Matching with Regular Expressions 17.5.2.1 The Syntax of Regular Expressions 17.5.2.2 Functions Using Regex 17.6 Dates with lubridate 17.6.1 ISO 8601 Format 17.6.2 Time-zones 17.6.3 Extract Date and Time Components 17.6.4 Calculating with Date-times 17.6.4.1 Durations 17.6.4.2 Periods 17.6.4.3 Intervals 17.6.4.4 Rounding 17.7 Factors with Forcats 18 Dealing with Missing Data 18.1 Reasons for Data to be Missing 18.2 Methods to Handle Missing Data 18.2.1 Alternative Solutions to Missing Data 18.2.2 Predictive Mean Matching (PMM) 18.3 R Packages to Deal with Missing Data 18.3.1 mice 18.3.2 missForest 18.3.3 Hmisc 19 Data Binning 19.1 What is Binning and Why Use It 19.2 Tuning the Binning Procedure 19.3 More Complex Cases: Matrix Binning 19.4 Weight of Evidence and Information Value 19.4.1 Weight of Evidence (WOE) 19.4.2 Information Value (IV) 19.4.3 WOE and IV in R 20 Factoring Analysis and Principle Components 20.1 Principle Components Analysis (PCA) 20.2 Factor Analysis V Modelling 21 Regression Models 21.1 Linear Regression 21.2 Multiple Linear Regression 21.2.1 Poisson Regression 21.2.2 Non-linear Regression 21.3 Performance of Regression Models 21.3.1 Mean Square Error (MSE) 21.3.2 R-Squared 21.3.3 Mean Average Deviation (MAD) 22 ClassificationModels 22.1 Logistic Regression 22.2 Performance of Binary Classification Models 22.2.1 The Confusion Matrix and Related Measures 22.2.2 ROC 22.2.3 The AUC 22.2.4 The Gini Coefficient 22.2.5 Kolmogorov-Smirnov (KS) for Logistic Regression 22.2.6 Finding an Optimal Cut-off 23 LearningMachines 23.1 Decision Tree 23.1.1 Essential Background 23.1.1.1 The Linear Additive Decision Tree 23.1.1.2 The CART Method 23.1.1.3 Tree Pruning 23.1.1.4 Classification Trees 23.1.1.5 Binary Classification Trees 23.1.2 Important Considerations 23.1.2.1 Broadening the Scope 23.1.2.2 Selected Issues 23.1.3 Growing Trees with the Package rpart 23.1.3.1 Getting Started with the Function rpart() 23.1.3.2 Example of a Classification Tree with rpart 23.1.3.3 Visualising a Decision Tree with rpart.plot 23.1.3.4 Example of a Regression Tree with rpart 23.1.4 Evaluating the Performance of a Decision Tree 23.1.4.1 The Performance of the Regression Tree 23.1.4.2 The Performance of the Classification Tree 23.2 Random Forest 23.3 Artificial Neural Networks (ANNs) 23.3.1 The Basics of ANNs in R 23.3.2 Neural Networks in R 23.3.3 The Work-flow to for Fitting a NN 23.3.4 Cross Validate the NN 23.4 Support Vector Machine 23.4.1 Fitting a SVM in R 23.4.2 Optimizing the SVM 23.5 Unsupervised Learning and Clustering 23.5.1 k-Means Clustering 23.5.1.1 k-Means Clustering in R 23.5.1.2 PCA before Clustering 23.5.1.3 On the Relation Between PCA and k-Means 23.5.2 Visualizing Clusters in Three Dimensions 23.5.3 Fuzzy Clustering 23.5.4 Hierarchical Clustering 23.5.5 Other Clustering Methods 24 Towards a Tidy Modelling Cycle with modelr 24.1 Adding Predictions 24.2 Adding Residuals 24.3 Bootstrapping Data 24.4 Other Functions of modelr 25 Model Validation 25.1 Model Quality Measures 25.2 Predictions and Residuals 25.3 Bootstrapping 25.3.1 Bootstrapping in Base R 25.3.2 Bootstrapping in the tidyverse with modelr 25.4 Cross-Validation 25.4.1 Elementary Cross Validation 25.4.2 Monte Carlo Cross Validation 25.4.3 k-Fold Cross Validation 25.4.4 Comparing Cross Validation Methods 25.5 Validation in a Broader Perspective 26 Labs 26.1 Financial Analysis with quantmod 26.1.1 The Basics of quantmod 26.1.2 Types of Data Available in quantmod 26.1.3 Plotting with quantmod 26.1.4 The quantmod Data Structure 26.1.4.1 Sub-setting by Time and Date 26.1.4.2 Switching Time Scales 26.1.4.3 Apply by Period 26.1.5 Support Functions Supplied by quantmod 26.1.6 Financial Modelling in quantmod 26.1.6.1 Financial Models in quantmod 26.1.6.2 A Simple Model with quantmod 26.1.6.3 Testing the Model Robustness 27 Multi Criteria Decision Analysis (MCDA) 27.1 What and Why 27.2 General Work-flow 27.3 Identify the Issue at Hand: Steps 1 and 2 27.4 Step 3: the Decision Matrix 27.4.1 Construct a Decision Matrix 27.4.2 Normalize the Decision Matrix 27.5 Step 4: Delete Inefficient and Unacceptable Alternatives 27.5.1 Unacceptable Alternatives 27.5.2 Dominance – Inefficient Alternatives 27.6 Plotting Preference Relationships 27.7 Step 5: MCDA Methods 27.7.1 Examples of Non-compensatory Methods 27.7.1.1 The MaxMin Method 27.7.1.2 The MaxMax Method 27.7.2 The Weighted Sum Method (WSM) 27.7.3 Weighted Product Method (WPM) 27.7.4 ELECTRE 27.7.4.1 ELECTRE I 27.7.4.2 ELECTRE II 27.7.4.3 Conclusions ELECTRE 27.7.5 PROMethEE 27.7.5.1 PROMethEE I 27.7.5.2 PROMethEE II 27.7.6 PCA (Gaia) 27.7.7 Outranking Methods 27.7.8 Goal Programming 27.8 Summary MCDA VI Introduction to Companies 28 Financial Accounting (FA) 28.1 The Statements of Accounts 28.1.1 Income Statement 28.1.2 Net Income: The P&L statement 28.1.3 Balance Sheet 28.2 The Value Chain 28.3 Further, Terminology 28.4 Selected Financial Ratios 29 Management Accounting 29.1 Introduction 29.1.1 Definition of Management Accounting (MA) 29.1.2 Management Information Systems (MIS) 29.2 Selected Methods in MA 29.2.1 Cost Accounting 29.2.2 Selected Cost Types 29.3 Selected Use Cases of MA 29.3.1 Balanced Scorecard 29.3.2 Key Performance Indicators (KPIs) 29.3.2.1 Lagging Indicators 29.3.2.2 Leading Indicators 29.3.2.3 Selected Useful KPIs 30 Asset Valuation Basics 30.1 Time Value of Money 30.1.1 Interest Basics 30.1.2 Specific Interest Rate Concepts 30.1.3 Discounting 30.2 Cash 30.3 Bonds 30.3.1 Features of a Bond 30.3.2 Valuation of Bonds 30.3.3 Duration 30.3.3.1 Macaulay Duration 30.3.3.2 Modified Duration 30.4 The Capital Asset Pricing Model (CAPM) 30.4.1 The CAPM Framework 30.4.2 The CAPM and Risk 30.4.3 Limitations and Shortcomings of the CAPM 30.5 Equities 30.5.1 Definition 30.5.2 Short History 30.5.3 Valuation of Equities 30.5.4 Absolute Value Models 30.5.4.1 Dividend Discount Model (DDM) 30.5.4.2 Free Cash Flow (FCF) 30.5.4.3 Discounted Cash Flow Model 30.5.4.4 Discounted Abnormal Operating Earnings Model 30.5.4.5 Net Asset Value Method or Cost Method 30.5.4.6 Excess Earnings Method 30.5.5 Relative Value Models 30.5.5.1 The Concept of Relative Value Models 30.5.5.2 The Price Earnings Ratio (PE) 30.5.5.3 Pitfalls when using PE Analysis 30.5.5.4 Other Company Value Ratios 30.5.6 Selection of Valuation Methods 30.5.7 Pitfalls in Company Valuation 30.5.7.1 Forecasting Performance 30.5.7.2 Results and Sensitivity 30.6 Forwards and Futures 30.7 Options 30.7.1 Definitions 30.7.2 Commercial Aspects 30.7.3 Short History 30.7.4 Valuation of Options at Maturity 30.7.4.1 A Long Call at Maturity 30.7.4.2 A Short Call at Maturity 30.7.4.3 Long and Short Put 30.7.4.4 The Put-Call Parity 30.7.5 The Black and Scholes Model 30.7.5.1 Pricing of Options Before Maturity 30.7.5.2 Apply the Black and Scholes Formula 30.7.5.3 The Limits of the Black and Scholes Model 30.7.6 The Binomial Model 30.7.6.1 Risk Neutral Method 30.7.6.2 The Equivalent Portfolio Binomial Model 30.7.6.3 Summary Binomial Model 30.7.7 Dependencies of the Option Price 30.7.7.1 Dependencies in a Long Call Option 30.7.7.2 Dependencies in a Long Put Option 30.7.7.3 Summary of Findings 30.7.8 The Greeks 30.7.9 Delta Hedging 30.7.10 Linear Option Strategies 30.7.10.1 Plotting a Portfolio of Options 30.7.10.2 Single Option Strategies 30.7.10.3 Composite Option Strategies 30.7.11 Integrated Option Strategies 30.7.11.1 The Covered Call 30.7.11.2 The Married Put 30.7.11.3 The Collar 30.7.12 Exotic Options 30.7.13 Capital Protected Structures VII Reporting 31 A Grammar of Graphics with ggplot2 31.1 The Basics of ggplot2 31.2 Over-plotting 31.3 Case Study for ggplot2 32 RMarkdown 33 knitr and LATEX 34 An Automated Development Cycle 35 Writing and Communication Skills 36 Interactive Apps 36.1 Shiny 36.2 Browser Born Data Visualization 36.2.1 HTML-widgets 36.2.2 Interactive Maps with leaflet 36.2.3 Interactive Data Visualisation with ggvis 36.2.3.1 Getting Started in R with ggvis 36.2.3.2 Combining the Power of ggvis and Shiny 36.2.4 googleVis 36.3 Dashboards 36.3.1 The Business Case: a Diversity Dashboard 36.3.2 A Dashboard with flexdashboard 36.3.2.1 A Static Dashboard 36.3.2.2 Interactive Dashboards with flexdashboard 36.3.3 A Dashboard with shinydashboard VIII Bigger and Faster R 37 Parallel Computing 37.1 Combine foreach and doParallel 37.2 Distribute Calculations over LAN with Snow 37.3 Using the GPU 37.3.1 Getting Started with gpuR 37.3.2 On the Importance of Memory use 37.3.3 Conclusions for GPU Programming 38 R and Big Data 38.1 Use a Powerful Server 38.1.1 Use R on a Server 38.1.2 Let the Database Server do the Heavy Lifting 38.2 Using more Memory than we have RAM 39 Parallelism for Big Data 39.1 Apache Hadoop 39.2 Apache Spark 39.2.1 Installing Spark 39.2.2 Running Spark 39.2.3 SparkR 39.2.3.1 A User Defined Function on Spark 39.2.3.3 Machine learning with SparkR 39.2.4 sparklyr 39.2.5 SparkR or sparklyr 40 The Need for Speed 40.1 Benchmarking 40.2 Optimize Code 40.2.1 Avoid Repeating the Same 40.2.2 Use Vectorisation where Appropriate 40.2.3 Pre-allocating Memory 40.2.4 Use the Fastest Function 40.2.5 Use the Fastest Package 40.2.6 Be Mindful about Details 40.2.7 Compile Functions 40.2.8 Use C or C++ Code in R 40.2.9 Using a C++ Source File in R 40.2.10 Call Compiled C++ Functions in R 40.3 Profiling Code 40.3.1 The Package profr 40.3.2 The Package proftools 40.4 Optimize Your Computer IX Appendices A Create your own R Package A.1 Creating the Package in the R Console A.2 Update the Package Description A.3 Documenting the Functions A.4 Loading the Package A.5 Further Steps B Levels ofMeasurement B.1 Nominal Scale B.2 Ordinal Scale B.3 Interval Scale B.4 Ratio Scale C Trademark Notices C.1 General Trademark Notices C.2 R-Related Notices C.2.1 Crediting Developers of R Packages C.2.2 The R-packages used in this Book D Code Not Shown in the Body of the Book E Answers to Selected Questions Bibliography Nomenclature Index
Similar books
MySQL® Notes for Professionals book
2018 · PDF
MrExcel 2022: Boosting Excel
2022 · PDF
MrExcel 2022: Boosting Excel
2022 · PDF
Session C11: Ancient Cultural Landscapes in South Europe – their Ecological Setting and Evolution, Session C22: Gardeners from South America, Session S04: Agro-Pastoralism and Early Metallurgy Sessions, Session WS29: The Idea of Enclosure in Recent Iberian Prehistory, Session C88: Rhytmes et causalites des dynamiques de l'anthropisation en Europe entre 6500 ET 500 BC: Hypotheses socio-culturelles et/ou climatiques: Proceedings of the XV UISPP World Congress (Lisbon 4-9 September 2006) / Actes du XV Congrès Mondial (Lisbonne 4-9 Septembre 2006) Vol.36
2010 · PDF
THE BRITISH ARMY IN INDIA: ITS PRESERVATION BY AN APPROPRIATE CLOTHING, HOUSING, LOCATING, RECREATIVE EMPLOYMENT, AND HOPEFUL ENCOURAGEMENT OF THE TROOPS. with AN APPENDIX ON INDIA : THE CLIMATE OP ITS HILLS ; THE DEVELOPMENT OF ITS RESODRCBS, INDUSTRY, AND ARTS ; THE ADMINISTRATION OF JUSTICE ; THE BLACK ACT ; THE PROGRESS OF CHRISTIANITY ; THE TRAFFIC IN OPIUM ; THE VALUE OF INDIA ; PERMANENT CAUSES OF DISAFFECTION, AND OF THE RECENT REBELLION ; THE TRADITIONARY POLICY; MISGOVERNMENT BY NATIVE RULERS ; ANNEXATIONS OF THEIR TERRITORY, ETC.
1858 · PDF
Idries Shah 27 Books Collection : A Perfumed Scorpion, A Veiled Gazelle, Caravan of Dreams, Darkest England, Destination Mecca, Evenings with Idries Shah, Knowing How to Know, Learning How to Learn, Letters and Lectures of Idries Shah, Neglected aspects of Sufi study, Observations, Oriental Magic, Reflections, Seeker after Truth, Special Illumination, Special Problems in the study of Sufi ideas, Sufi thought and action, Tales of the Dervishes, The Dermis Probe, The Elephant in the Dark, The Englishman Handbook, Idries Shah Antology, The Magic Monastery, The natives are restless, wisdom of the Idiots PDF.
2022 · PDF
The travels of Capts. Lewis and Clarke from St. Louis, by way of the Missouri and Columbia rivers, to the Pacific ocean; performed in the years 1804, 1805 & 1806, by order of the government of the United States. Containing delineations of the manners, customs, religion, &c. of the Indians, comp. from various authentic sources, and original documents, and a summary of the Statistical view of the Indian nations, from the official communication of Meriwether Lewis. Illustrated with a map of the country, inhabited by the western tribes of Indians
1809 · PDF
Professional Linux kernel architecture ''Wrox programmer to programmer''--Cover. - ''What you are reading right now is the result of an evolution over more than seven years: After two years of writing, the first edition was published in German by Carl Hanser Verlag in 2003. It then described kernel 2.6.0. The test was used as a basis for the low-level design documentation for the EAL4+ security evaluation of Red Hat Enterprise Linux 5, requiring to update it to kernel 2.6.18 (if the EAL acronym does not mean anything to you, then Wikipedia is once more your friend). Hewlett-Packard sponsored the translation into English and has, thankfully, granted the rights to publish the result. Updates to kernel 2.6.24 were then performed specifically for this book''--P. ix
2008 · PDF