Reliable Machine Learning: Applying SRE Principles to ML in Production. Early Release
Book information
Description
Whether you are part of a small startup or a planet-spanning megacorp, this practical book shows data scientists, SREs, and business owners how to run ML reliably, effectively, and accountably within your organization. You'll gain insight into everything from how to do model monitoring in production to how to run a well-tuned model development team in a product organization. By applying an SRE mindset to machine learning, authors and engineering professionals Cathy Chen, Kranti Parisa, Niall Richard Murphy, D. Sculley, Todd Underwood, and featured guests show you how to run an efficient ML system. Whether you want to increase revenue, optimize decision-making, solve problems, or understand and influence customer behavior, you'll learn how to perform day-to-day ML tasks while keeping the bigger picture in mind. You'll examine: What ML is: how it functions and what it relies on Conceptual frameworks for understanding how ML "loops" work Effective "productionization," and how it can be made easily monitorable, deployable, and operable Why ML systems make production troubleshooting more difficult, and how to get around them How ML, product, and production teams can communicate effectively Prospective Table of Contents (Subject to Change) Preface Why We Wrote This Book SRE as the lens on ML Intended Audience How this book is organized Our Approach Let’s Knit! Navigating this Book About the Authors Acknowledgments 1. Introduction The ML Lifecycle Data collection and analysis ML Training Pipelines Build/Integrate and Validate Applications Launch Monitoring and Feedback Loops Lessons from the Loop 2. Data Management Principles Data as Liability The Data Sensitivity of ML Pipelines Phases of Data Creation Ingestion Processing Storage Management Analysis & Visualization Data Reliability Durability Consistency Version Control Performance Availability Data Integrity Security Privacy Policy & Compliance In Conclusion and Next Steps Policy and Governance Data Sciences and Software Infrastructure Infrastructure 3. Basic Introduction to Models What is a model? A Basic Model Creation Workflow Model Architecture vs. Configured Model vs. Trained Model Where Are the Vulnerabilities? Training Data Labels Training Methods Infrastructure and Pipelines Platforms Feature Generation Upgrades and Fixes A Set of Useful Questions to Ask about Any Model An Example ML System Yarn Product Click Prediction Model Features Labels for Features Model Updating Model Serving Common Failures In Conclusion 4. Feature and Training Data Features Feature Selection and Engineering Lifecycle of a feature Feature Systems Feature Store Feature Quality Evaluation System Labels Human Generated Labels Annotation Workforces Measuring Human Annotation Quality An Annotation Platform Active Learning and AI Assisted Labeling Documentation and Training for Labelers Metadata Metadata Systems Overview Dataset Metadata Feature Metadata Label Metadata Pipeline Metadata Data Privacy and Fairness Privacy Fairness Conclusion 5. Evaluating Model Validity and Quality Model Validity Evaluating Model Quality Offline Evaluations Operationalizing Verification and Evaluation 6. Fairness, Privacy, and Ethical Machine Learning Systems Fairness (a.k.a. Fighting bias) Definitions of fairness Reaching fairness Fairness as a process rather than an endpoint A quick legal note Privacy Methods to preserve privacy A quick legal note Responsible AI Explanation Effectiveness Social and Cultural Appropriateness Responsible AI along the ML Pipeline Use case brainstorming Data collection and cleaning Model creation and training Model validation and quality assessment Model deployment Products for the market In Conclusion 7. Training Systems Requirements Basic Training System Implementation Features Feature Storage Model Management System Orchestration Job/Process/Resource Scheduling System ML framework Model quality evaluation system Monitoring General Reliability Principles Most failures will not be ML failures Models will be retrained Models will have multiple versions (at the same time!) Good models will become bad Data will be unavailable Models should be improvable Features will be added and changed Models can train too fast Resource Utilization Matters Utilization != Efficiency Outages include recovery Common Training Reliability Problems Data sensitivity Reproducibility Example Reproducibility problem at YarnIt Compute Resource Capacity Example Capacity problem at YarnIt Structural Reliability Organizational challenges Ethics and fairness considerations Conclusion 8. Serving Key Questions for Model Serving What will be the load to our model? What are the prediction latency needs of our model? Where does the model need to live? What are the hardware needs for our model? How will the serving model be stored, loaded, versioned, and updated? What will our feature pipeline for serving look like? Model Serving Architectures Offline Serving (Batch Inference) Advantages: Disadvantages: Online Serving (Online Inference) Advantages: Disadvantages: Model as a Service (MaaS) Advantages: Disadvantages: Serving at the Edge Advantages: Disadvantages: Choosing an architecture Model API Design Testing Serving for Accuracy or Resilience? Scaling Autoscaling Caching Disaster Recovery Ethics and fairness considerations Conclusions 9. Monitoring and Observability for Models What is production monitoring, and why do it? What does it look like? The concerns that machine learning brings to monitoring Reasons for continual ML observability—in production Problems with ML production monitoring Difficulties of development versus serving Sidebar: an important note about skew A mindset change is required Best practices for ML model monitoring Generic Pre-Serving Model Recommendations Training and retraining Model validation (before rollout)13 Serving Conclusion 10. Continuous ML Anatomy of a Continuous ML System Observations About Continuous ML Systems External world events may influence our systems Models can influence their own training data Temporal effects can arise at several different time scales Emergency response must be done in real time New launches require staged ramp-ups and stable baselines Models must be managed rather than shipped. Continuous Organizations Rethinking Non-Continuous ML Systems 11. Incident Response Incident Management Basics Life of an Incident Incident Response Roles Anatomy of an ML-centric Outage Terminology Reminder: “Model” Story Time Story 1: Searching But Not Finding Story 2: Suddenly Useless Partners Story 3: Recommend You Find New Suppliers ML Incident Management Principles Guiding Principles Special Topics Production Engineers and ML Engineering vs Modeling The Ethical On-Call Engineer Manifesto Conclusion 12. How Product and ML Interact Different Types of Products Agile ML? ML Product Development Phases Discovery and Definition Setting Business Goals Build a Minimal Viable Product (MVP) Model and Product Development Production Deployment Support & Maintenance Build vs Buy Models Data Processing Infrastructure End-to-End Platforms Scoring approach for making the decision Making the decision Sample yarnit.ai store features powered by ML Showcasing popular yarns by total sales Recommendations based on browsing history Cross-selling and upselling Content-based filtering Collaborative filtering Conclusion 13. Integrating ML Into Your Organization Chapter Assumptions Leader-based viewpoint Detail Matters ML needs to know about the business The most important assumption you make The value of ML Significant organizational risks ML is not magic Mental (Way of Thinking) Model Inertia Surfacing risk correctly in different cultures Siloed teams don’t solve all problems Implementation models Remembering the goal Greenfield versus Brownfield ML roles and responsibilities How to hire ML folks Organizational Design and Incentives Strategy Structure Processes Rewards People A note on sequencing Conclusion 14. Practical ML Org Implementation Examples Scenario 1: A New Centralized ML team Background and Organizational Description Process Rewards People Default Implementation Scenario 2: Decentralized ML infrastructure and expertise Background and Organization Description Process Rewards People Default Implementation Scenario 3: Hybrid: Centralized Infrastructure/Decentralized Modeling Background and Organization Description Rewards People Default Implementation Conclusion 15. Case Studies: ML Ops in Practice 1. Respecting and dealing with Privacy and Data Retention policies in ML pipelines Background Challenge #1: Which dialects? Challenge #2: Racing with time Takeaways 2. Continuous ML model impacting traffic Background Problem & Resolution Takeaways 3. Steel Inspection Background Problem & Resolution Takeaways 4. NLP MLOps—Profiling and staging load test Background Problem & Resolution Takeaways 5. Ad Click Prediction: Databases Versus Reality Background Problem & Resolution Takeaways 6. Testing and measuring dependencies in ML workflow Background Takeaways About the Authors
Similar books
Reliable Machine Learning: Applying SRE Principles to ML in Production
2022 · EPUB
Reliable Machine Learning
2022 · EPUB
MySQL® Notes for Professionals book
2018 · PDF
MrExcel 2022: Boosting Excel
2022 · PDF
MrExcel 2022: Boosting Excel
2022 · PDF
Session C11: Ancient Cultural Landscapes in South Europe – their Ecological Setting and Evolution, Session C22: Gardeners from South America, Session S04: Agro-Pastoralism and Early Metallurgy Sessions, Session WS29: The Idea of Enclosure in Recent Iberian Prehistory, Session C88: Rhytmes et causalites des dynamiques de l'anthropisation en Europe entre 6500 ET 500 BC: Hypotheses socio-culturelles et/ou climatiques: Proceedings of the XV UISPP World Congress (Lisbon 4-9 September 2006) / Actes du XV Congrès Mondial (Lisbonne 4-9 Septembre 2006) Vol.36
2010 · PDF
THE BRITISH ARMY IN INDIA: ITS PRESERVATION BY AN APPROPRIATE CLOTHING, HOUSING, LOCATING, RECREATIVE EMPLOYMENT, AND HOPEFUL ENCOURAGEMENT OF THE TROOPS. with AN APPENDIX ON INDIA : THE CLIMATE OP ITS HILLS ; THE DEVELOPMENT OF ITS RESODRCBS, INDUSTRY, AND ARTS ; THE ADMINISTRATION OF JUSTICE ; THE BLACK ACT ; THE PROGRESS OF CHRISTIANITY ; THE TRAFFIC IN OPIUM ; THE VALUE OF INDIA ; PERMANENT CAUSES OF DISAFFECTION, AND OF THE RECENT REBELLION ; THE TRADITIONARY POLICY; MISGOVERNMENT BY NATIVE RULERS ; ANNEXATIONS OF THEIR TERRITORY, ETC.
1858 · PDF
Idries Shah 27 Books Collection : A Perfumed Scorpion, A Veiled Gazelle, Caravan of Dreams, Darkest England, Destination Mecca, Evenings with Idries Shah, Knowing How to Know, Learning How to Learn, Letters and Lectures of Idries Shah, Neglected aspects of Sufi study, Observations, Oriental Magic, Reflections, Seeker after Truth, Special Illumination, Special Problems in the study of Sufi ideas, Sufi thought and action, Tales of the Dervishes, The Dermis Probe, The Elephant in the Dark, The Englishman Handbook, Idries Shah Antology, The Magic Monastery, The natives are restless, wisdom of the Idiots PDF.
2022 · PDF