ENGLISH

Fundamentals of Predictive Analytics With Jmp

Book information

Publisher
SAS Institute
Year
2016
ISBN
1629598569, 9781629598567
Language
english
Format
PDF
Filesize
18 MB (19280062 bytes)
Pages
406\406
Time added
2020-04-03 19:50:11

Description

Written for students in undergraduate and graduate statistics courses, as well as for the practitioner who wants to make better decisions from data and models, this updated and expanded second edition of Fundamentals of Predictive Analytics with JMP(r) bridges the gap between courses on basic statistics, which focus on univariate and bivariate analysis, and courses on data mining and predictive analytics. Going beyond the theoretical foundation, this book gives you the technical knowledge and problem-solving skills that you need to perform real-world multivariate data analysis. First, this book teaches you to recognize when it is appropriate to use a tool, what variables and data are required, and what the results might be. Second, it teaches you how to interpret the results and then, step-by-step, how and where to perform and evaluate the analysis in JMP(r). Using JMP(r) 13 and JMP(r) 13 Pro, this book offers the following new and enhanced features in an example-driven format: an add-in for Microsoft Excel Graph Builder dirty data visualization regression ANOVA logistic regression principal component analysis LASSO elastic net cluster analysis decision trees k-nearest neighbors neural networks bootstrap forests boosted trees text mining association rules model comparison With today's emphasis on business intelligence, business analytics, and predictive analytics, this second edition is invaluable to anyone who needs to expand his or her knowledge of statistics and to apply real-world, problem-solving analysis. This book is part of the SAS Press progr Contents About The Book What Does This Book Cover? Is This Book for You? What’s New in This Edition? What Should You Know about the Examples? Software Used to Develop the Book's Content Example Code and Data Where Are the Exercise Solutions? We Want to Hear from You About These Authors Acknowledgments Introduction Historical Perspective Two Questions Organizations Need to Ask Return on Investment Cultural Change Business Intelligence and Business Analytics Figure 1.1: A Framework of Business Analytics Introductory Statistics Courses Figure 1.2: A Student’s View of a Statistical Study from a Basic Statistics Course The Problem of Dirty Data Added Complexities in Multivariate Analysis Practical Statistical Study Obtaining and Cleaning the Data Figure 1.3: The Flow of a Real-World Statistical Study Understanding the Statistical Study as a Story The Plan-Perform-Analyze-Reflect Cycle Figure 1.4: The PPAR Cycle Using Powerful Software Framework and Chapter Sequence Figure 1.5: A Framework for Multivariate Analysis Statistics Review Introduction Fundamental Concepts 1 and 2 FC1: Always Take a Random and Representative Sample Figure 2.1: Population Distribution of the Weights of Sumo Wrestlers and Jockeys Figure 2.2: Population and a Sample Distribution of the Weights of Sumo Wrestlers and Jockeys FC2: Remember That Statistics Is Not an Exact Science Fundamental Concept 3: Understand a Z-Score Fundamental Concept 4 FC4: Understand the Central Limit Theorem Figure 2.3: Population Distribution and Sample Distribution of Observations and Sampling Distribution of the Means for the Weights of Sumo Wrestlers and Jockeys Learn from an Example Figure 2.4: Excel Data Analysis Tool Histogram Dialog Box Figure 2.5: Results of the Histogram Data Analysis Tool Figure 2.6: Histogram of a Random Sample of 30 Sumo Wrestler and Jockeys Weights Fundamental Concept 5 Understand One-Sample Hypothesis Testing Consider p-Values Table 2.1: Decisions and Conclusions to Hypothesis Tests in Relationship to the p-Value Fundamental Concept 6: Understand That Few Approaches/Techniques Are Correct—Many Are Wrong Illustration 1 Table 2.2 Data and Descriptive Statistics in Countif.xls file and Worksheet Statistics Ways JMP Can Access Data in Excel Figure 2.7: Excel Import Wizard Dialog Box Figure 2.8: Modeling Types of Gender Figure 2.9: The Data Table for Countif.jmp after Modeling Type Changes Figure 2.10: The JMP Distribution Dialog Box Figure 2.11: Distribution Output for Countif.jmp Data Illustration 2 Figure 2.12: Fit Y by X Dialog Box Figure 2.13: Bivariate Analysis of Salary by Years Figure 2.14: Contingency Analysis of Major by Gender Three Possible Outcomes When You Choose a Technique Dirty Data Introduction Figure 3.1: A Framework to Multivariate Analysis Data Set Error Detection Figure 3.2: Distribution Dialog Box Figure 3.3: Distribution Output for Variables Default, Reason, and Job Figure 3.4: Data Table for hmeq.jmp Figure 3.5: Recode Dialog Box Outlier Detection Approach 1 Figure 3.6: Distribution Output for Variables Loan, Mortgage, and Value Figure 3.7: Data Table of Outliers for Variable Value Approach 2 Figure 3.8: Quantile Range Outliers Report for Variable Value Table 3.1: Outlier Values for Variable Value and the Suggested Corrections Missing Values Table 3.2: An Example of Patterns of Missingness for 100 Observations (0 Implies Missing) Statistical Assumptions of Patterns of Missing Missing Completely at Random Missing at Random Missing Not at Random Conventional Correction Methods Listwise Deletion Method Table 3.2: A Small Example of Missing Data Variable Removal Method Conventional Estimate Value Methods Mean, Median, or Mode Substitution Dummy Variable Regression Imputation Model-Based Methods Maximum Likelihood Methods Multiple Imputation Methods The JMP Approach Example Using JMP Figure 3.9: Missing Data Pattern Dialog Box Figure 3.10: Missing Data Pattern Figure 3.11 New Data Table after Creating Indicator Columns Figure 3.12: Explore Missing Values Report Figure 3.13: Explore Missing Values Report with Only Continuous Variables General First Steps on Receipt of a Data Set Exercises Data Discovery with Multivariate Data Introduction Figure 4.1: A Conceptual View of a Data Set Figure 4.2: A Framework for Multivariate Analysis Use Tables to Explore Multivariate Data PivotTables Figure 4.3 PivotTable Worksheet Figure 4.4: Value Field Setting Dialog Box Figure 4.5: Resulting Excel PivotTable Tabulate in JMP Figure 4.6: The Tabulate Dialog Box Figure 4.7: Selection Button to Copy the Output Figure 4.8: Resulting Copy of the Tabulate Table and Chart (into a Microsoft Word Document) Use Graphs to Explore Multivariate Data Graph Builder Figure 4.9: Graph Builder Dialog Box Figure 4.10 Graph of Salary and Year When Year Is in Drop Zone Group X Figure 4.11 Graph of Salary and Year When Year Is in Drop Zone Group Y Figure 4.12 Graph of Salary and Year When Year Is in Drop Zone Wrap Figure 4.13 Graph of Salary and Year When Year Is in Drop Zone Overlap Figure 4.14: Salary versus Year, by Major and Gender Scatterplot Figure 4.15: Scatterplot Matrix of Countif.jmp Figure 4.16: Scatterplot with Shaded Ellipses Explore a Larger Data Set Trellis Chart Figure 4.17: Trellis Chart of the Home Price Indices, by State Figure 4.18: Trellis Chart of Home Price Index and CPI by State Bubble Plot Figure 4.19 Bubble Plot Dialog Box Figure 4.20: Bubble Plot of Home Price Index by State on 3/1975 Figure 4.21: Bubble Plot of Home Price Index by State on 3/2005 Figure 4.22: Bubble Plot of Home Price Index by State on 3/2008 Figure 4.23: Bubble Plot of Home Price Index by State on 9/2009 Explore a Real-World Data Set Use Graph Builder to Examine Results of Analyses Figure 4.24: Revenue by Quarter Using Contours Figure 4.25: Revenue by Quarter Using Line Graphs Generate a Trellis Chart and Examine Results Figure 4.26: Trellis Chart of the Average Revenue, Cost of Sales, and Gross Product over Time (Quarter) by Product Line Figure 4.27: Graph of the Average Revenue, Cost of Sales, and Gross Product, by Quarter for the Credit Products Product Line Figure 4.28: Graph of the Average Revenue, Cost of Sales, and Gross Product, by Quarter for the Credit Products Product Line, by Customer ID Figure 4.29: Bar Chart of the Average Revenue and %GP of Credit Products Customers Use Dynamic Linking to Explore Comparisons in a Small Data Subset Return to Graph Builder to Sort and Visualize a Larger Data Set Figure 4.30: Bar Chart of the Average Revenue and %GP of Credit Products Customers in Ascending Order Figure 4.31: Two Data Filters with Scroll bar Figure 4.32: Bar Chart Using Further Elements of the Data Filter Regression and ANOVA Introduction Figure 5.1: A Framework for Multivariate Analysis Regression Perform a Simple Regression and Examine Results Figure 5.2: Scatterplot with Corresponding Simple Linear Regression Understand and Perform Multiple Regression Examine the Correlations and Scatterplot Matrix Figure 5.3: The Multivariate and Correlations Dialog Box Figure 5.4: Correlations and Scatterplot Matrix for the salesperfdata.jmp Data Perform a Multiple Regression Figure 5.5: The Fit Model Dialog Box Examine the Effect Summary Report and Adjust the Model Figure 5.6: Multiple Linear Regression Output (Top Half) Figure 5.7: Multiple Linear Regression Output (Bottom Half) Figure 5.8: Parameter Estimates Table with Standardized Betas and Variance Inflation Factors Evaluate the Statistical Significance of the Model Step 1a: Conduct an F Test Step 1b: Evaluate the T Test for Each Independent Variable Step 1c: Examine the Residual Plot Figure 5.9: Residual Plot of Predicted versus Residual Step 1d: Assess Multicollinearity Step 2: Determine the Goodness of Fit Figure 5.10: Stepwise Regression for the Sales Performance Data Set Understand and Perform Regression with Categorical Data Figure 5.11: Regression of Age on Number of Words Memorized Figure 5.12: Formula Dialog Box Figure 5.13: Regression of Age1 on Number of Words Memorized Table 5.1: Dummy Variable Coding for Process and the Predicted Number of Words Memorized Figure 5.14: Regression of Process on Number of Words Memorized Analysis of Variance Perform a One-Way ANOVA Figure 5.15: One-Way ANOVA of Age and Words Evaluate the Model Test Statistical Assumptions Figure 5.16: Tests of Unequal Variances and Welch’s Test for Age and Words Create a Normal Quantile Plot Figure 5.17: Normal Quantile Plot of the Residuals Perform a One-Factor ANOVA Figure 5.18: One-Way ANOVA of Process and Words Examine the Results Figure 5.19: Tests of Unequal Variances and Welch’s Test for Process and Words Test for Differences Figure 5.20: One-Way ANOVA of Gender and Pain Figure 5.21: One-Way ANOVA of Drug and Pain Student’s t Test Tukey Honest Significant Difference Test Figure 5.22: The Tukey-Kramer HSD Test Hsu’s Multiple Comparison with Best Figure 5.23: The Hsu’s MCB Test Dunnett’s Test The JMP Fit Model Platform Figure 5.24: One-Way ANOVA Output Using Fit Model Perform a Two-Way ANOVA Examine the Results Figure 5.25: Two-way ANOVA Output (top half) Figure 5.26: Two-way ANOVA Output (bottom half) Figure 5.27: Two-way ANOVA Output for Gender and Drug Figure 5.28: Two-Way ANOVA Output for Gender * Drug Evaluate the Model for Equal Variances Figure 5.29: Formula Dialog Box Figure 5.30: One-Way ANOVA with Pain and Interaction Gender * Drug Exercises Logistic Regression Introduction Dependence Technique Figure 6.1: A Framework for Multivariate Analysis The Linear Probability Model The Logistic Function Figure 6.2: The Logistic Function A Straightforward Example Using JMP Create a Dummy Variable Figure 6.3: Formula Dialog Box Use a Contingency Table to Determine the Odds Ratio Figure 6.4: Control Panel for Tabulate Figure 6.5: Contingency Table from toydataset.jmp Calculate the Odds Ratio Method 1: Compute the Probabilities Method 2: Run a Logistic Regression Figure 6.6: Ranges of Probabilities, Odds, and Log-odds Figure 6.7: Fit Model Dialog Box Figure 6.8: Logistic Regression Results for toylogistic.jmp Examine the Parameter Estimates Figure 6.9: Odds Ratios Tables Using the Nominal Independent Variable PassMidterm Change the Default Convention Figure 6.10: Changing the Value Order Examine the Changed Results Figure 6.11: Parameter Estimates Figure 6.12: Odds Ratios Tables Using the Continuous Independent Variable MidtermScore Compute Probabilities for Each Observation Figure 6.13: Verifying Calculation of Probability of Failing Figure 6.14: Confusion Matrix Check the Model Figure 6.15: Scatterplot of Lin[0] and MidtermScore Figure 6.16: Whole Model Test for the Toylogistic Data Set Figure 6.17: Lack-of-Fit-Test for Current Model A Realistic Logistic Regression Statistical Study Understand the Model-Building Approach Figure 6.18: Distribution of Intl_Calls and VMail_Message Run Bivariate Analyses Run the Initial Regression and Examine the Results Figure 6.19: Whole Model Test and Lack of Fit for the Churn Data Set Convert a Continuous Variable to Discrete Variables Figure 6.20: Histogram of CustServ_Call Figure 6.21: Creating the CustServ Variable Produce Interaction Variables Figure 6.22: Logistic Regression Results with Interaction Term Added Validate and Use the Model Figure 6.23: Confusion Matrix Exercises Principal Components Analysis Introduction Figure 7.1: A Framework for Multivariate Analysis Basic Steps in JMP Produce the Correlations and Scatterplot Matrix Figure 7.2: Correlations and Scatterplot Matrix for the toyprincomp.xls Data Set Create the Principal Components Figure 7.3: Principal Components Dialog Box Figure 7.4: Principal Components Summary Plots Run a Regression of y on Prin1 and Excluding Prin2 Table 7.1: Regression Results of y on Prin1 and Excluding Prin2 Understand Eigenvalue Analysis Conduct the Eigenvalue Analysis and the Bartlett Test Figure 7.5: Eigenvectors and Eigenvalues for the toyprincomp.jmpData Set Verify Lack of Correlation Dimension Reduction Produce the Correlations and Scatterplot Matrix Conduct the Principal Component Analysis Determine the Number of Principal Components to Select Figure 7.6: Scree Plot and Eigenvalues for x1 to x12 from the princomp.jmp Data Set Method 1 Method 2 Method 3 Compare Methods for Determining the Number of Components Table 7.2: Regression Results Discovery of Structure in the Data A Straightforward Example Produce the Correlations and Scatterplot Matrix Conduct the Principal Component Analysis Figure 7.7: PCA Summary Plots, Eigenvalues, and Scree Plot for the olymp88sas.jmp Data Set An Example with Less Well Defined Data Conduct a Principal Component Analysis and Correct Mistakes Rerun the Principal Component Analysis Figure 7.8: PCA Summary Plots for the StateGDP2008.jmp Data Set Refine Results by Expressing Data in Proportional Form Figure 7.9: PCA Summary Plots for the StateGDP2008percent.jmp Data Set Exercises Least Absolute Shrinkage and Selection Operator and Elastic Net Introduction Figure 8.1: A Framework for Multivariate Analysis The Importance of the Bias-Variance Tradeoff Figure 8.2: Two Estimation Distributions Ridge Regression Technique and Limitations Use of JMP Table 8.1: Hypothetical Values of Ridge Criterion for Various Values of λ Figure 8.3: The Generalized Regression Dialog Box Figure 8.4: Generalized Regression Output for the Mass Housing Data Set Least Absolute Shrinkage and Selection Operator Perform the Technique Examine the Results Figure 8.5: LASSO Regression for the Mass Housing Data Set Refine the Results Elastic Net Perform the Technique Examine the Results Figure 8.6: Elastic Net Regression for the Mass Housing Data Set Compare with LASSO Exercises Cluster Analysis Introduction Figure 9.1: A Framework for Multivariate Analysis Example Applications An Example from the Credit Card Industry The Need to Understand Statistics and the Business Problem Hierarchical Clustering Understand the Dendrogram Understand the Methods for Calculating Distance between Clusters Figure 9.2: Ways to Measure Distance between Two Clusters Figure 9.3: The Different Effects of Single and Complete Linkage Perform a Hierarchal Clustering with Complete Linkage Table 9.1: Toy Data Set for Illustrating Hierarchical Clustering Figure 9.4: The Clustering Dialog Box Figure 9.5: Dendrogram of the Toy Data Set Examine the Results Figure 9.6: Scatterplot of the Toy Data Set Consider a Scree Plot to Discern the Best Number of Clusters Apply the Principles to a Small but Rich Data Set Table 9.2: Thompson’s 1975 Public Utility Data Set Figure 9.7: Hierarchical Clustering of the Public Utility Data Set Figure 9.8: Cluster Means for Five Clusters Consider Adding Clusters in a Regression Analysis K-Means Clustering Understand the Benefits and Drawbacks of the Method Figure 9.9: The Function of k-Means Clustering Choose k and Determine the Clusters Figure 9.10: U-Shaped Sum of Squared Errors Plot for Choosing the Number of Clusters Figure 9.11: A Scree Plot for Choosing the Number of Clusters Perform k-Means Clustering Figure 9.12: Determining How Many Clusters Are in the Data Set Change the Number of Clusters Figure 9.13: KMeans Dialog Box Figure 9.14: Output of Means Clustering with Five Clusters Figure 9.15: Output of k Means Clustering with Three Clusters Create a Profile of the Clusters with Parallel Coordinate Plots Figure 9.16: Parallel Coordinate Plots for k = 3 Clusters Figure 9.17: Formula Editor for Creating Distance Squared Perform Iterative Clustering Table 9.3: Five Clusters and Their Members for the Public Utility Data Figure 9.18: Cluster Means for Five Clusters Using k-Means Score New Observations K-Means Clustering versus Hierarchical Clustering Exercises Decision Trees Introduction Figure 10.1: A Framework for Multivariate Analysis Benefits and Drawbacks Definitions and an Example Figure 10.2: Classifying Bank Customers as “Good” or “Bad” Risks for a Loan Theoretical Questions Classification Trees Table 10.1: The Variables in the freshmen1.jmp Data Set Begin Tree and Observe Results Figure 10.3: Partition Initial Output with Discrete Dependent Variable Use JMP to Choose the Split That Maximizes the LogWorth Statistic Figure 10.4: Initial Rate, Probabilities, and LogWorths Split the Root Node According to Rank of Variables Figure 10.5: Decision Tree after First Split Split Second Node According to the College Variable Figure 10.6: Fit Model Details Figure 10.7: Candidate Variables for Splitting an Impure Node Figure 10.8: Splitting a Node (n = 89) Examine Results and Predict the Variable for a Third Split Figure 10.9: Splitting a Node (n = 27) Examine Results and Predict the Variable for a Fourth Split Figure 10.10: Splitting a Node (n = 22) Examine Results and Continue Splitting to Gain Actionable Insights Prune to Simplify Overgrown Trees Examine Receiver Operator Characteristic and Lift Curves Figure 10.11: Receiver Operator Characteristic and Lift Curves Regression Trees Figure 10.12: Partition Initial Output with Continuous Discrete Dependent Variable Understand How Regression Trees Work Figure 10.13: RSquare Report Table before Splitting Figure 10.14: Table Containing Statistics from Several Splits Figure 10.15: Plot of AIC and RMSE by Number of Splits Restart a Regression Driven by Practical Questions Use Column Contributions and Leaf Reports for Large Data Sets Figure 10.16: Column Contributions Figure 10.17: Leaf Report Exercises k-Nearest Neighbors Introduction Figure 11.1: A Framework for Multivariate Analysis Example—Age and Income as Correlates of Purchase Figure 11.2: Whether Purchased or Not, by Age and Income Table 11.1: Customer Data and Distance from Customer E The Way That JMP Resolves Ties The Need to Standardize Units of Measurement Table 11.2: Customer Raw Data and Distance from Customer E Table 11.3: Customer Standardized Data k-Nearest Neighbors Analysis Perform the Analysis Figure 11.3: k-Nearest Neighbor Output for toylogistic.jmp File Make Predictions for New Data Figure 11.4: Example of the Nonsymmetrical Nature of the k-Nearest Neighbor Algorithm k-Nearest Neighbor for Multiclass Problems Understand the Variables Table 11.4: Types of Glass Table 11.5: Ten Attributes of Glass Perform the Analysis and Examine Results Figure 11.5: k-Nearest Neighbor for the glass.jmp File The k-Nearest Neighbor Regression Models Perform a Linear Regression as a Basis for Comparison Apply the k-Nearest Neighbors Technique Figure 11.6: k-Nearest Neighbor Results for the MassHousing.jmp File Compare the Two Methods Figure 11.7: Overlap Plot for the MassHousing.jmp File Figure 11.8: Residual Plot for Multiple Linear Regression Make Predictions for New Data Limitations and Drawbacks of the Technique Exercises Neural Networks Introduction Figure 12.1: A Framework for Multivariate Analysis Drawbacks and Benefits A Simplified Representation Figure 12.2: A Neuron Accepting Weighted Inputs Figure 12.3: A Neuron with a Bias Term Figure 12.4: Hyperbolic Tangent Activation Function A More Realistic Representation Figure 12.5: A Standard Neural Network Architecture Understand Validation Methods Holdback Validation Figure 12.6: Typical Error Based on the Training Sample and the Holdback Sample k-fold Cross-Validation Understand the Hidden Layer Structure A Few Guidelines for Determining Number of Nodes Practical Strategies for Determining Number of Nodes The Method of Boosting Understand Options for Improving the Fit of a Model Complete the Data Preparation Table 12.1: Converting an Ordered Categorical Variable to [0,1] Use JMP on an Example Data Set Perform a Linear Regression as a Baseline Figure 12.7: Neural Network Model Launch Figure 12.8: Results of Neural Network Using Default Options Perform the Neural Network Ten Times to Assess Default Performance Table 12.2 Training and Validation RSquare Running the Default Model Ten Times Boost the Default Model Table 12.3 Training and Validation RSquare When Boosting the Default Model, Number of Models = 100 Compare Transformation of Variables and Methods of Validation Table 12.4: R2 for Five Runs of Neural Networks Figure 12.9: Residual Plot for Training Data When R2 = 74% Figure 12.10: Residual Plot for Validation Data When R2 = 88% Table 12.5: R2 for Different Penalty Functions Table 12.6: R2 for Various Architectures Figure 12.11: Creating a Binary Variable for Median Price Figure 12.12: Default Model for the Binary Dependent Variable, MedPrice Exercises Bootstrap Forests and Boosted Trees Introduction Figure 13.1: A Framework for Multivariate Analysis Bootstrap Forests Understand Bagged Trees Perform a Bootstrap Forest Table 13.1: Variables in the TitanicPassengers.jmp Data Set Understand the Options in the Dialog Box Figure 13.2: The Bootstrap Forest Dialog Box Figure 13.3: Bootstrap Forest Output for the Titanic Passengers Data Set Select Options and Relaunch Figure 13.4: Bootstrap Forest Output with the Number of Terms Sampled per Split to 2 Examine the Improved Results Perform a Bootstrap Forest for Regression Trees Figure 13.5: Bootstrap Forest Output for the Mass Housing Data Set Boosted Trees Understand Boosting Perform Boosting Figure 13.6: The Boosted Tree Dialog Box Understand the Options in the Dialog Box Select Options and Relaunch Figure 13.7: Boosted Tree Output for the Titanic Passengers Data Set Examine the Improved Results Figure 13.8: Boosted Tree Output with a Learning Rate of 0.9 Perform a Boosted Tree for Regression Trees Figure 13.9: Boosted Tree Output for the Mass Housing Data Set Use Validation and Training Samples Create a Dummy Variable Figure 13.10: The Make Validation Column Report Dialog Box Perform a Boosting at Default Settings Examine Results and Relaunch Figure 13.11: Boosted Trees Results for the Titanic Passengers Data Set with a Training and Validation Set Compare Results to Choose the Least Misleading Model Figure 13.12: Boosted Trees Results with Learning Rate of 0.9 Figure 13.13: Boosted Trees Results with and Learning Rate of 0.4 Exercises Model Comparison Introduction Figure 14.1: A Framework for Multivariate Analysis Perform a Model Comparison with Continuous Dependent Variable Understand Absolute Measures Understand Relative Measures Understand Correlation between Variable and Prediction Explore the Uses of the Different Measures Table 14.1: Performance Measures for the McDonalds48.jmp File Figure 14.2: Scatterplot with Out-of-Sample Predictions as Plus Signs Perform a Model Comparison with Binary Dependent Variable Understand the Confusion Matrix and Its Limitations Table 14.2: An Error Table or Confusion Matrix Understand True Positive Rate and False Positive Rate Interpret Receiving Operator Characteristic Curves Figure 14.3: An ROC Curve Figure 14.4: ROC Curves and Line of Optimal Classification Compare Two Example Models Predicting Churn Figure 14.5: ROC Curves for Logistic (Left) and Partition (Right) Figure 14.6: Data Filter Perform a Model Comparison Using the Lift Chart Table 14.3: Lift Values Figure 14.7: Initial Lift Curves for Logistic (Left) and Classification Tree (Right) Figure 14.8: Lift Curves for Logistic (Left) and Classification Tree (Right) Train, Validate, and Test Perform Stepwise Regression Figure 14.9: Control Panel for Stepwise Regression Figure 14.10: Regression Output for Model Chosen by Stepwise Regression Examine the Results of Stepwise Regression Compute the MSE, MAE, and Correlation Table 14.4: Performance Measures for the McDonalds72.jmp File Examine the Results for MSE, MAE, and Correlation Understand Overfitting from a Coin-Flip Example Table 14.5: The Number of Heads Observed When Each Coin Was Tossed 50 Times Table 14.6: The Number of Heads in 50 Tosses with the Three Coins That You Believe to Be Biased Use the Model Comparison Platform Continuous Dependent Variable Perform the Stepwise Regression Perform the Linear Regression Figure 14.11: The Generalized Regression Output for the Mass Housing Data Set with Nonzero Variables Perform the LASSO Figure 14.12: Solution Path Graphs and Parameter Estimates for Original Predictors Compare the Models Using the Training and Validation Sets Perform Forward Selection and LASSO on the Training Set Figure 14.13: Forward Selection Generalized Regression for Training Data from the Mass Housing Data Set Figure 14.14: Regression Output for Mass Housing Data Set, with Forward Selection Perform LASSO on the Validation Set Compare Model Results for Training and Validation Observations Figure 14.15: Regression Output for Mass Housing Data Set with Lasso Figure 14.16: Model Comparison Dialog Box Figure 14.17: Model Comparison of Mass Housing Data Set Using Forward Selection and Lasso Discrete Dependent Variable Figure 14.18: Model Comparison Dialog Box Figure 14.19: Model Comparison Output Using the Churn Data Set Exercises Text Mining Introduction Historical Perspective Unstructured Data Figure 15.1: A Framework for Multivariate Analysis Developing the Document Term Matrix Figure 15.2: Flowchart of the Stages of Text Processing Figure 15.3: Data Table of toytext.jmp File Understand the Tokenizing Stage Figure 15.4: Text Explorer Dialog Box Select Options in the Text Explorer Dialog Box Figure 15.5: Text Explorer Output Box Figure 15.6: Toytext.jmp Data Table with Initial Document Terms Recode to Correct Misspellings and Group Terms Figure 15.7: Recode Dialog Box Figure 15.8: The Recode Dialog Box Figure 15.9: Text Explorer Output Box Figure 15.10: Text Explorer Output Box after Stemming Understand the Phrasing Stage Figure 15.11: Text Explorer Output Box after Phrasing Figure 15.12: Document Text Matrix Understand the Terming Stage Create Stop Words Figure 15.13: Text Explorer Output Box after Terming Generate a Word Cloud Observe the Order of Operations Developing the Document Term Matrix with a Larger Data Set Generate a Word Cloud and Examine the Text Figure 15.14: Text Explorer Output Box Examine and Group Terms Add Frequent Phrases to List of Terms Figure 15.15: Text Explorer Output Box Parse the List of Terms Figure 15.16: Text Explorer Output Box Using Multivariate Techniques Perform Latent Semantic Analysis Understanding SVD Matrices Plot the Documents or Terms Figure 15.17: Latent Semantic Analysis Specifications Dialog Box Figure 15.18: SVD Plots Figure 15.19: Top Portion of SVD Scatterplots of SVD Plots of 10 Singular Vectors Perform Topic Analysis Figure 15.20: Topic Analysis Output Perform Cluster Analysis Begin the Analysis Examine the Results Figure 15.21: Top Portion of Latent Class Analysis for 10 Clusters Output Box Figure 15.22: Lower Portion of Latent Class Analysis for 10 Clusters Output Box Figure 15.23: Upper Leftmost Portion of SVD Scatterplots Identify Dominant Terms Using Predictive Techniques Figure 15.24: Singular Values Perform Primary Analysis Figure 15.25: Distribution Output for Violation Type Perform Logistic Regressions Figure 15.26: Top Portion of Initial Logistic Regression Output Figure 15.27: Lower Portion of Initial Logistic Regression Output Figure 15.28: Top Portion with Singular Values Logistic Regression Output Figure 15.29: Lower Portion with Singular Values Logistic Regression Output Exercises Market Basket Analysis Introduction Association Analyses Figure 16.1: Framework for Multivariate Analysis Examples Understand Support, Confidence, and Lift Table 16.1: Five Small Market Baskets Association Rules Support Confidence Lift Use JMP to Calculate Confidence and Lift Use the A Priori Algorithm for More Complex Data Sets Form Rules and Calculate Confidence and Lift Analyze a Real Data Set Perform Association Analysis with Default Settings Reduce the Number of Rules and Sort Them Examine Results Figure 16.2: Confidence versus Lift for the GroceriesPurchase Data Figure 16.3: Frequent Item Sets Target Results to Take Business Actions Exercises Statistical Storytelling The Path from Multivariate Data to the Modeling Process Early Applications of Data Mining Credit Card Fraud Detection Customer Relationship Management Numerous JMP Customer Stories of Modern Applications Definitions of Data Mining Data Mining Predictive Analytics A Framework for Predictive Analytics Techniques Figure 17.1: A Framework for Predictive Analytics Techniques The Goal, Tasks, and Phases of Predictive Analytics The Difference between Statistics and Data Mining Table 17.1: The Data Mining Process and the Percentage of Time Spent on Each Phase SEMMA Index A B C D E F G H I J K L M N O P R S T U V W Z

Similar books

Session C11: Ancient Cultural Landscapes in South Europe – their Ecological Setting and Evolution, Session C22: Gardeners from South America, Session S04: Agro-Pastoralism and Early Metallurgy Sessions, Session WS29: The Idea of Enclosure in Recent Iberian Prehistory, Session C88: Rhytmes et causalites des dynamiques de l'anthropisation en Europe entre 6500 ET 500 BC: Hypotheses socio-culturelles et/ou climatiques: Proceedings of the XV UISPP World Congress (Lisbon 4-9 September 2006) / Actes du XV Congrès Mondial (Lisbonne 4-9 Septembre 2006) Vol.36

Session C11: Ancient Cultural Landscapes in South Europe – their Ecological Setting and Evolution, Session C22: Gardeners from South America, Session S04: Agro-Pastoralism and Early Metallurgy Sessions, Session WS29: The Idea of Enclosure in Recent Iberian Prehistory, Session C88: Rhytmes et causalites des dynamiques de l'anthropisation en Europe entre 6500 ET 500 BC: Hypotheses socio-culturelles et/ou climatiques: Proceedings of the XV UISPP World Congress (Lisbon 4-9 September 2006) / Actes du XV Congrès Mondial (Lisbonne 4-9 Septembre 2006) Vol.36

2010 · PDF

THE BRITISH ARMY IN INDIA: ITS PRESERVATION BY AN APPROPRIATE CLOTHING, HOUSING, LOCATING, RECREATIVE EMPLOYMENT, AND HOPEFUL ENCOURAGEMENT OF THE TROOPS. with AN APPENDIX ON INDIA : THE CLIMATE OP ITS HILLS ; THE DEVELOPMENT OF ITS RESODRCBS, INDUSTRY, AND ARTS ; THE ADMINISTRATION OF JUSTICE ; THE BLACK ACT ; THE PROGRESS OF CHRISTIANITY ; THE TRAFFIC IN OPIUM ; THE VALUE OF INDIA ; PERMANENT CAUSES OF DISAFFECTION, AND OF THE RECENT REBELLION ; THE TRADITIONARY POLICY; MISGOVERNMENT BY NATIVE RULERS ; ANNEXATIONS OF THEIR TERRITORY, ETC.

THE BRITISH ARMY IN INDIA: ITS PRESERVATION BY AN APPROPRIATE CLOTHING, HOUSING, LOCATING, RECREATIVE EMPLOYMENT, AND HOPEFUL ENCOURAGEMENT OF THE TROOPS. with AN APPENDIX ON INDIA : THE CLIMATE OP ITS HILLS ; THE DEVELOPMENT OF ITS RESODRCBS, INDUSTRY, AND ARTS ; THE ADMINISTRATION OF JUSTICE ; THE BLACK ACT ; THE PROGRESS OF CHRISTIANITY ; THE TRAFFIC IN OPIUM ; THE VALUE OF INDIA ; PERMANENT CAUSES OF DISAFFECTION, AND OF THE RECENT REBELLION ; THE TRADITIONARY POLICY; MISGOVERNMENT BY NATIVE RULERS ; ANNEXATIONS OF THEIR TERRITORY, ETC.

1858 · PDF

Idries Shah 27 Books Collection : A Perfumed Scorpion, A Veiled Gazelle, Caravan of Dreams, Darkest England, Destination Mecca, Evenings with Idries Shah, Knowing How to Know, Learning How to Learn, Letters and Lectures of Idries Shah, Neglected aspects of Sufi study, Observations, Oriental Magic, Reflections, Seeker after Truth, Special Illumination, Special Problems in the study of Sufi ideas, Sufi thought and action, Tales of the Dervishes, The Dermis Probe, The Elephant in the Dark, The Englishman Handbook, Idries Shah Antology, The Magic Monastery, The natives are restless, wisdom of the Idiots PDF.

Idries Shah 27 Books Collection : A Perfumed Scorpion, A Veiled Gazelle, Caravan of Dreams, Darkest England, Destination Mecca, Evenings with Idries Shah, Knowing How to Know, Learning How to Learn, Letters and Lectures of Idries Shah, Neglected aspects of Sufi study, Observations, Oriental Magic, Reflections, Seeker after Truth, Special Illumination, Special Problems in the study of Sufi ideas, Sufi thought and action, Tales of the Dervishes, The Dermis Probe, The Elephant in the Dark, The Englishman Handbook, Idries Shah Antology, The Magic Monastery, The natives are restless, wisdom of the Idiots PDF.

2022 · PDF

The travels of Capts. Lewis and Clarke from St. Louis, by way of the Missouri and Columbia rivers, to the Pacific ocean; performed in the years 1804, 1805 & 1806, by order of the government of the United States. Containing delineations of the manners, customs, religion, &c. of the Indians, comp. from various authentic sources, and original documents, and a summary of the Statistical view of the Indian nations, from the official communication of Meriwether Lewis. Illustrated with a map of the country, inhabited by the western tribes of Indians

The travels of Capts. Lewis and Clarke from St. Louis, by way of the Missouri and Columbia rivers, to the Pacific ocean; performed in the years 1804, 1805 & 1806, by order of the government of the United States. Containing delineations of the manners, customs, religion, &c. of the Indians, comp. from various authentic sources, and original documents, and a summary of the Statistical view of the Indian nations, from the official communication of Meriwether Lewis. Illustrated with a map of the country, inhabited by the western tribes of Indians

1809 · PDF

Professional Linux kernel architecture ''Wrox programmer to programmer''--Cover. - ''What you are reading right now is the result of an evolution over more than seven years: After two years of writing, the first edition was published in German by Carl Hanser Verlag in 2003. It then described kernel 2.6.0. The test was used as a basis for the low-level design documentation for the EAL4+ security evaluation of Red Hat Enterprise Linux 5, requiring to update it to kernel 2.6.18 (if the EAL acronym does not mean anything to you, then Wikipedia is once more your friend). Hewlett-Packard sponsored the translation into English and has, thankfully, granted the rights to publish the result. Updates to kernel 2.6.24 were then performed specifically for this book''--P. ix

Professional Linux kernel architecture ''Wrox programmer to programmer''--Cover. - ''What you are reading right now is the result of an evolution over more than seven years: After two years of writing, the first edition was published in German by Carl Hanser Verlag in 2003. It then described kernel 2.6.0. The test was used as a basis for the low-level design documentation for the EAL4+ security evaluation of Red Hat Enterprise Linux 5, requiring to update it to kernel 2.6.18 (if the EAL acronym does not mean anything to you, then Wikipedia is once more your friend). Hewlett-Packard sponsored the translation into English and has, thankfully, granted the rights to publish the result. Updates to kernel 2.6.24 were then performed specifically for this book''--P. ix

2008 · PDF