ENGLISH

Machine Learning Engineering

Book information

Year
2020
ISBN
9781777005450
Language
english
Format
PDF
Filesize
38 MB (39872259 bytes)
Pages
\310
Time added
2021-07-11 18:32:45

Description

Foreword Preface Who This Book is For How to Use This Book Should You Buy This Book? Introduction Notation and Definitions Data Structures Capital Sigma Notation What is Machine Learning Supervised Learning Unsupervised Learning Semi-Supervised Learning Reinforcement Learning Data and Machine Learning Terminology Data Used Directly and Indirectly Raw and Tidy Data Training and Holdout Sets Baseline Machine Learning Pipeline Parameters vs. Hyperparameters Classification vs. Regression Model-Based vs. Instance-Based Learning Shallow vs. Deep Learning Training vs. Scoring When to Use Machine Learning When the Problem Is Too Complex for Coding When the Problem Is Constantly Changing When It Is a Perceptive Problem When It Is an Unstudied Phenomenon When the Problem Has a Simple Objective When It Is Cost-Effective When Not to Use Machine Learning What is Machine Learning Engineering Machine Learning Project Life Cycle Summary Before the Project Starts Prioritization of Machine Learning Projects Impact of Machine Learning Cost of Machine Learning Estimating Complexity of a Machine Learning Project The Unknowns Simplifying the Problem Nonlinear Progress Defining the Goal of a Machine Learning Project What a Model Can Do Properties of a Successful Model Structuring a Machine Learning Team Two Cultures Members of a Machine Learning Team Why Machine Learning Projects Fail Lack of Experienced Talent Lack of Support by the Leadership Missing Data Infrastructure Data Labeling Challenge Siloed Organizations and Lack of Collaboration Technically Infeasible Projects Lack of Alignment Between Technical and Business Teams Summary Data Collection and Preparation Questions About the Data Is the Data Accessible? Is the Data Sizeable? Is the Data Useable? Is the Data Understandable? Is the Data Reliable? Common Problems With Data High Cost Bad Quality Noise Bias Low Predictive Power Outdated Examples Outliers Data Leakage What Is Good Data Good Data Is Informative Good Data Has Good Coverage Good Data Reflects Real Inputs Good Data Is Unbiased Good Data Is Not a Result of a Feedback Loop Good Data Has Consistent Labels Good Data Is Big Enough Summary of Good Data Dealing With Interaction Data Causes of Data Leakage Target is a Function of a Feature Feature Hides the Target Feature From the Future Data Partitioning Leakage During Partitioning Dealing with Missing Attributes Data Imputation Techniques Leakage During Imputation Data Augmentation Data Augmentation for Images Data Augmentation for Text Dealing With Imbalanced Data Oversampling Undersampling Hybrid Strategies Data Sampling Strategies Simple Random Sampling Systematic Sampling Stratified Sampling Storing Data Data Formats Data Storage Levels Data Versioning Documentation and Metadata Data Lifecycle Data Manipulation Best Practices Reproducibility Data First, Algorithm Second Summary Feature Engineering Why Engineer Features How to Engineer Features Feature Engineering for Text Why Bag-of-Words Works Converting Categorical Features to Numbers Feature Hashing Topic Modeling Features for Time-Series Use Your Creativity Stacking Features Stacking Feature Vectors Stacking Individual Features Properties of Good Features High Predictive Power Fast Computability Reliability Uncorrelatedness Other Properties Feature Selection Cutting the Long Tail Boruta L1-Regularization Task-Specific Feature Selection Synthesizing Features Feature Discretization Synthesizing Features from Relational Data Synthesizing Features from the Data Synthesizing Features from Other Features Learning Features from Data Word Embeddings Document Embeddings Embeddings of Anything Choosing Embedding Dimensionality Dimensionality Reduction Fast Dimensionality Reduction with PCA Dimensionality Reduction for Visualization Scaling Features Normalization Standardization Data Leakage in Feature Engineering Possible Problems Solution Storing and Documenting Features Schema File Feature Store Feature Engineering Best Practices Generate Many Simple Features Reuse Legacy Systems Use IDs as Features when Needed… …But Reduce the Cardinality When Possible Use Counts with Caution Make Feature Selection When Necessary Test the Code Carefully Keep Code, Model, and Data in Sync Isolate Feature Extraction Code Serialize Together Model and Feature Extractor Log the Values of Features Summary Supervised Model Training (Part 1) Before You Start Working on the Model Validate Schema Conformity Define an Achievable Performance Level Choose a Performance Metric Choose the Right Baseline Split Data Into Three Sets Preconditions for Supervised Learning Representing Labels for Machine Learning Multiclass Classification Multi-label Classification Selecting the Learning Algorithm Main Properties of a Learning Algorithm Algorithm Spot-Checking Building a Pipeline Assessing Model Performance Performance Metrics for Regression Performance Metrics for Classification Performance Metrics for Ranking Hyperparameter Tuning Grid Search Random Search Coarse-to-Fine Search Other Techniques Cross-Validation Shallow Model Training Shallow Model Training Strategy Saving and Restoring the Model Bias-Variance Tradeoff Underfitting Overfitting The Tradeoff Regularization L1 and L2 Regularization Other Forms of Regularization Summary Supervised Model Training (Part 2) Deep Model Training Strategy Neural Network Training Strategy Performance Metric and Cost Function Parameter-Initialization Strategies Optimization Algorithms Learning Rate Decay Schedules Regularization Network Size Search and Hyperparameter Tuning Handling Multiple Inputs Handling Multiple Outputs Transfer Learning Stacking Models Types of Ensemble Learning An Algorithm of Model Stacking Data Leakage in Model Stacking Dealing With Distribution Shift Types of Distribution Shift Adversarial Validation Handling Imbalanced Datasets Class Weighting Ensemble of Resampled Datasets Other Techniques Model Calibration Well-Calibrated Models Calibration Techniques Troubleshooting and Error Analysis Reasons for Poor Model Behavior Iterative Model Refinement Error Analysis Error Analysis in Complex Systems Using Sliced Metrics Fixing Wrong Labels Finding Additional Examples to Label Troubleshooting Deep Learning Best Practices Deliver a Good Model Trust Popular Open Source Implementations Optimize a Business-Specific Performance Measure Upgrade From Scratch Avoid Correction Cascades Use Model Cascading With Caution Write Efficient Code, Compile, and Parallelize Test on Both Newer and Older Data More Data Beats Cleverer Algorithm New Data Beats Cleverer Features Embrace Tiny Progress Facilitate Reproducibility Summary Model Evaluation Offline and Online Evaluation A/B Testing G-Test Z-Test Concluding Remarks and Warnings Multi-Armed Bandit Statistical Bounds on the Model Performance Statistical Interval for the Classification Error Bootstrapping Statistical Interval Bootstrapping Prediction Interval for Regression Evaluation of Test Set Adequacy Neuron Coverage Mutation Testing Evaluation of Model Properties Robustness Fairness Summary Model Deployment Static Deployment Dynamic Deployment on User's Device Deployment of Model Parameters Deployment of a Serialized Object Deploying to Browser Advantages and Drawbacks Dynamic Deployment on a Server Deployment on a Virtual Machine Deployment in a Container Serverless Deployment Model Streaming Deployment Strategies Single Deployment Silent Deployment Canary Deployment Multi-Armed Bandits Automated Deployment, Versioning, and Metadata Model Accompanying Assets Version Sync Model Version Metadata Model Deployment Best Practices Algorithmic Efficiency Deployment of Deep Models Caching Delivery Format for Model and Code Start With a Simple Model Test on Outsiders Summary Model Serving, Monitoring, and Maintenance Properties of the Model Serving Runtime Security and Correctness Ease of Deployment Guarantees of Model Validity Ease of Recovery Avoidance of Training/Serving Skew Avoidance of Hidden Feedback Loops Modes of Model Serving Serving in Batch Mode Serving on Demand to a Human Serving on Demand to a Machine Model Serving in Real World Being Ready for Errors Dealing With Errors Being Ready for, and Dealing With, Change Being Ready for, and Dealing With, Human Nature Model Monitoring What Can Go Wrong? What and How to Monitor What to Log Monitor for Abuse Model Maintenance When to Update How to Update Summary Conclusion Takeaways What to Read Next Acknowledgements Index

Similar books