ENGLISH

Becoming a Rockstar SRE: Electrify your site reliability engineering mindset to build reliable, resilient, and efficient systems

Book information

Publisher
Packt Publishing
Year
2023
ISBN
1803239220, 9781803239224
Language
english
Format
PDF
Filesize
31 MB (32534223 bytes)
Pages
420\420
Time added
2023-06-14 00:41:30

Description

Excel in site reliability engineering by learning from field-driven lessons on observability and reliability in code, architecture, process, systems management, costs, and people to minimize downtime and enhance developers' output Purchase of the print or Kindle book includes a free eBook in the PDF format Key FeaturesUnderstand the goals of an SRE in terms of reliability, efficiency, and constant improvementMaster highly resilient architecture in server, serverless, and containerized workloadsLearn the why and when of employing Kubernetes, GitHub, Prometheus, Grafana, Terraform, Python, Argo CD, and GitOpsBook Description Site reliability engineering is all about continuous improvement, finding the balance between business and product demands while working within technological limitations to drive higher revenue. But quantifying and understanding reliability, handling resources, and meeting developer requirements can sometimes be overwhelming. With a focus on reliability from an infrastructure and coding perspective, Becoming a Rockstar SRE brings forth the site reliability engineer (SRE) persona using real-world examples. This book will acquaint you the role of an SRE, followed by the why and how of site reliability engineering. It walks you through the jobs of an SRE, from the automation of CI/CD pipelines and reducing toil to reliability best practices. You'll learn what creates bad code and how to circumvent it with reliable design and patterns. The book also guides you through interacting and negotiating with businesses and vendors on various technical matters and exploring observability, outages, and why and how to craft an excellent runbook. Finally, you'll learn how to elevate your site reliability engineering career, including certifications and interview tips and questions. By the end of this book, you'll be able to identify and measure reliability, reduce downtime, troubleshoot outages, and enhance productivity to become a true rockstar SRE! What you will learnGet insights into the SRE role and its evolution, starting from Google's original visionUnderstand the key terms, such as golden signals, SLO, SLI, MTBF, MTTR, and MTTDOvercome the challenges in adopting site reliability engineeringEmploy reliable architecture and deployments with serverless, containerization, and release strategiesIdentify monitoring targets and determine observability strategyReduce toil and leverage root cause analysis to enhance efficiency and reliabilityRealize how business decisions can impact quality and reliabilityWho this book is for This book is for IT professionals, including developers looking to advance into an SRE role, system administrators mastering technologies, and executives experiencing repeated downtime in their organizations. Anyone interested in bringing reliability and automation to their organization to drive down customer impact and revenue loss while increasing development throughput will find this book useful. A basic understanding of API and web architecture and some experience with cloud computing and services will assist with understanding the concepts covered. Table of ContentsSRE Job Role - Activities and ResponsibilitiesFundamental Numbers - Reliability StatisticsImperfect Habits - Duct Tape Architecture and Spaghetti CodeEssential Observability - Metrics, Events, Logs, and Traces (MELT)Resolution Path - Master TroubleshootingOperational Framework - Managing Infrastructure and SystemsData Consumed - Observability Data Science (N.B. Please use the Look Inside option to see further chapters) Cover Title Page Copyright and Credits Dedication Contributors Table of Contents Preface Part 1 - Understanding the Basics of Who, What, and Why Chapter 1: SRE Job Role – Activities and Responsibilities Making this journey personal SRE driving forces SRE skills SRE traits Understanding the mindset and hobbies of an SRE SRE affinity game SRE guiding principles SRE hobbies DevOps engineers versus SRE versus others DevOps and site reliability engineers Software and site reliability engineers Describing an SRE’s main responsibilities An overview of the daily activities of an SRE People that inspire Jeremy’s recognition – Paul Tyma, former CTO, LendingTree Rod’s recognition – Ingo Averdunk, Distinguished Engineer, IBM, and Gene Brown, Distinguished Engineer, Kyndryl Summary Further reading Chapter 2: Fundamental Numbers – Reliability Statistics SLA commitment – a conversation, not a number Internal partner SLAs External partner SLAs The cost of more 9s in an SLA A final word on SLAs Defining and leveraging SLOs and SLIs SLOs SLOs and time Tracking outage frequency with the MTBF Measuring the downtime with the MTTR Understanding the customer and revenue impact Transparency in outages The rockstar SRE’s SLA Summary Chapter 3: Imperfect Habits – Duct Tape Architecture and Spaghetti Code The business of software development – let’s start with the dollars Defining the “value” of software to a business The value of protecting business The value of growing a business The value of saving labor costs The A/B testing mindset – the art of change in customer interaction A/B testing in customer flows Analyzing the results of A/B testing Leveraging A/B testing to satisfy quarterly numbers Dedication to the craft of development – and why some are just here for a job A quick guide to communicating with your colleagues Reviewing the merge request – it’s about training, oversight, and reliability Avoiding the typical rubber stamp mentality A word on production deployments Why businesses want us to outright ignore best practices The truth about the ownership of a developer’s time Understanding the flaws in how we estimate development cost Fast, good, cheap – pick one Why is observability the answer to reliability issues? The cost of highly available architecture Mixing good and bad – tricks to wrapping bad code and making it resilient Alerting that fires actions Adding additional logging to monitor potential issues Using try catch to encapsulate exceptions Retries to the rescue…or not Summary Part 2 - Implementing Observability for Site Reliability Engineering Chapter 4: Essential Observability – Metrics, Events, Logs, and Traces (MELT) Technical requirements Accomplishing systems monitoring and telemetry Monitoring targets for infrastructure Monitoring types and tools Monitoring golden signals Monitoring data Understanding APM Getting to know topology self-discovery, the blast radius, predictability, and correlation Alerting – the art of doing it quietly The user perspective notification trigger principle Event-to-incident mapping principle Mixing everything into observability Outages versus downtime Observability architecture Observability effectiveness In practice – applying what you have learned Lab architecture Lab contents Lab instructions Summary Further reading Chapter 5: Resolution Path – Master Troubleshooting Properly defining the problem – and what to ask and not ask Source of information The knowledge base of the reporter Naming conventions False urgency Executive summary Breaking down and testing systems Breaking down hardware versus the operating system Breaking down a web API Understanding the steps The problems with this method of troubleshooting Previous and common events – checking for the simple problems Prior Root Cause Analysis (RCA) documents Timeline analysis Comparison The best approach Effective research both online and among peers The art of the Google search Skimming the content quickly and refining it Never forget your internal resources Breaking down source code efficiently Code you’ve never seen When that fails Logging plus code In practice – applying what you’ve learned Summary Chapter 6: Operational Framework – Managing Infrastructure and Systems Technical requirements Approaching systems administration as a discipline Design Installation Configuration App deployment Management Upgrade Uninstallation Understanding IT service management ITIL DevOps Seeing systems administration as multiple layers and multiple towers Automating systems provisioning and management Infrastructure as Code Immutable infrastructure In practice – applying what you’ve learned Lab architecture Lab contents Lab instructions Summary Further readings Chapter 7: Data Consumed – Observability Data Science Technical requirements Making data-driven decisions Defining the question and options Determining which data to use Identifying which data is already available Collecting the missing data Analyzing all datasets together Presenting the decision as a record Documenting the lessons learned in the process Solving problems through a scientific approach Formulation Hypothesis Prediction Experiment Analysis Understanding the most common statistical methods Percentages Mean, average, and standard deviation Quantiles and percentiles Histograms Using other mathematical models in observability Visualizing histograms with Grafana In practice – applying what you’ve learned Lab architecture Lab contents Lab instructions Summary Further reading Part 3 - Applying Architecture for Reliability Chapter 8: Reliable Architecture – Systems Strategy and Design Technical requirements Designing for reliability Architectural aspects Reliability equations Design patterns Modern applications Splitting and balancing the workload Splitting Balancing Failing over – almost as good Scaling up and out – horizontal versus vertical Horizontal Vertical Autoscaling In practice – applying what you’ve learned Lab architecture Lab contents Lab instructions Summary Further reading Chapter 9: Valued Automation – Toil Discovery and Elimination Technical requirements Eliminating toil Toil redefined Why toil is bad Handling toil the right way Treating automation as a software problem Document Algorithm Code Automating the (in)famous CI/CD pipeline Continuous integration Continuous delivery Production releases In practice – applying what you’ve learned Lab architecture Lab contents Lab instructions Summary Further reading Chapter 10: Exposing Pipelines – GitOps and Testing Essentials A basic pipeline – building automation to deploy infrastructure as code architecture and code Pipelines in chronological order Pipeline templates Errors or breaks in pipelines Using containers in pipelines Pipeline artifacts Pipeline troubleshooting tips Automating compliance and security in pipelines Library age Application security testing Dynamic Application Security Testing (DAST) Static Application Security Testing (SAST) Secrets scanning Automated linting for code quality and standards Compiling with linting feedback Validating functionality during deployment with automated testing Why is testing so important to reliability? Test data The types of testing When to test a pipeline Testing observability Automated rollbacks The reduction of developer toil through automated processes What is the impact of addressing toil? In practice – applying what you’ve learned Preparing AWS for the lab Creating your repository Adding secrets to your repository Downloading and committing the lab files Understanding the pipeline Adding more steps Testing but not deploying Lab final thoughts Summary Chapter 11: Worker Bees – Orchestrations of Serverless, Containers, and Kubernetes Technical requirements The multiple definitions of serverless Serverless Framework Serverless computing Serverless functions Monitoring serverless functions Errors Containers and why we love them Isolation Immutability Promotability Tagging Rollbacks Security Signable Monitoring containers Kubernetes and other ways to orchestrate containers Health checks Crashing and force-closing containers HTTP-based load balancing Server load balancing Containers as a Service (CaaS) Simple container orchestration Kubernetes Deployment techniques and workers Traditional replacement deployment Rolling deployment A/B or blue/green deployment Canary deployment Automation and rolling back failed deployments Rollback metrics When to roll back How to roll back In practice – applying what you’ve learned Leveraging Gitpod – a containerized workspace The emulation source code Running the emulation Summary Chapter 12: Final Exam – Tests and Capacity Planning Technical requirements Understanding types of testing Development tests Build tests Delivery tests Deployment tests Production tests Adopting TDD Unit testing the hard way Unit testing with a framework Using test automation frameworks Staying ahead with capacity planning Load test data The capacity curve The demand curve In practice – applying what you’ve learned Lab architecture Lab contents Lab instructions Summary Further reading Part 4 - Mastering the Outage Moments Chapter 13: First Thing – Runbooks and Low Noise Outage Notifications Technical requirements What makes a good runbook – the basics Runbooks as living documents Understanding the runbook audience knowledge level Runbook audience permissions What do you put into a runbook anyway? Beyond the runbook – code and comments Quickly understanding source code Searching source code for your needle in a haystack Commenting for understanding What’s in a good dashboard? Types of dashboards NOC-style red and green Displaying trends Aggregates and breakdowns What dashboards are not The basics of priority levels Response effort Engineer retention Incident response systems and priority Incident response systems and phone-based alerts What is a priority one event? Defining priority based on... The priority level of observability failures Forcing the priority – the rockstar way! Adjusting alerts Logs and alerting Pausing alerts In practice – applying what you’ve learned Defining priority levels Custom hat pricing API runbook Alerting Summary Chapter 14: Rapid Response – Outage Management Techniques Where to meet – an effective strategy for communicating good information Online collaboration In-person collaboration The historical data found in outage responses Participants Follow-up work Leveraging the people involved in the response Tasks Participants and personalities Break strategy and stress management The opportunity to respond at the right time Training Runbook and contact list revisions Team building Executive messaging bugs in the ear Opportunities to call out during the RCA Messaging customers and leadership Customer versus leadership messaging Cadence Email groups Status sites Over-messaging Notes, notes, notes... In practice – applying what you’ve learned Outage and alarm Notification and response Troubleshooting The conclusion Summary Chapter 15: Postmortem Candor – Long-Term Resolution The content of the postmortem in executive summary style Executive summary style Overview Impact Timeline Detailed technical description Response Resolution Future actions Decisions are not blame Business is business Resource and time constraints Monitoring The cost of more reliability as a business decision Active:Active Manual failover Cost of time to identify The cost of time to move a load Hidden development costs Training and skill sets – they matter Identifying gaps Training and certification targets Creating future action plans Immediate follow-up Who to involve Timelines and priority Assigning ownership Tracking the work In-practice – an example of a postmortem Writing the overview Rounding out the postmortem Custom Hat Company postmortem Impact Timeline Technical details and response Resolution Future actions Summary Part 5 - Looking into Future Trends and Preparing for SRE Interviews Chapter 16: Chaos Injector – Advanced Systems Stability Technical requirements Comprehending the wheel-of-misfortune game All ends are new beginnings Lessons to be learned Role-playing scenarios A little bit of gamification Understanding chaos engineering for reliability Principles of chaos engineering Chaos system architecture Chaos experiments In practice – employing the wheel-of-misfortune game Lab architecture Lab contents Lab instructions In practice – injecting chaos into systems Lab architecture Lab contents Lab instructions Summary Further reading Chapter 17: Interview Advice – Hiring and Being Hired What we’re looking for in a candidate Are you qualified? Entry-level SRE job Problem-solving The ability to accept feedback and direction A broad knowledge base and skill set Research and learning skill set The ability to say “No” Culture fit The X factor Passion Experience Personal responsibility Common interview questions and answers Technical questions Non-technical questions Insightfully odd questions What should you look for in a career? Define a good boss Dotted line reporting Morals Researching the company Business model Profitability for the next decade Structure Large versus small Public versus private Online reviews Are you over-or under-certified? Certifications that matter How many are too many certifications? Relevancy Tips for landing the job with a great salary Interview tips Salary negotiations Summary Appendix A The Site Reliability Engineer Manifesto The manifesto How to adopt it How to contribute to it Appendix B The 12-Factor App Questionnaire The questionnaire Factor I – Code base Factor II – Dependencies Factor III – Config (configuration) Factor IV – Backing (backend) services Factor V – Build, release, run Factor VI – Processes Factor VII – Port binding Factor VIII – Concurrency Factor IX – Disposability Factor X – Development/production (dev/prod) parity Factor XI – Logs Factor XII – Admin processes How to adopt this questionnaire How to contribute to this questionnaire Index About Packt Other Books You May Enjoy

Similar books

Session C11: Ancient Cultural Landscapes in South Europe – their Ecological Setting and Evolution, Session C22: Gardeners from South America, Session S04: Agro-Pastoralism and Early Metallurgy Sessions, Session WS29: The Idea of Enclosure in Recent Iberian Prehistory, Session C88: Rhytmes et causalites des dynamiques de l'anthropisation en Europe entre 6500 ET 500 BC: Hypotheses socio-culturelles et/ou climatiques: Proceedings of the XV UISPP World Congress (Lisbon 4-9 September 2006) / Actes du XV Congrès Mondial (Lisbonne 4-9 Septembre 2006) Vol.36

Session C11: Ancient Cultural Landscapes in South Europe – their Ecological Setting and Evolution, Session C22: Gardeners from South America, Session S04: Agro-Pastoralism and Early Metallurgy Sessions, Session WS29: The Idea of Enclosure in Recent Iberian Prehistory, Session C88: Rhytmes et causalites des dynamiques de l'anthropisation en Europe entre 6500 ET 500 BC: Hypotheses socio-culturelles et/ou climatiques: Proceedings of the XV UISPP World Congress (Lisbon 4-9 September 2006) / Actes du XV Congrès Mondial (Lisbonne 4-9 Septembre 2006) Vol.36

2010 · PDF

THE BRITISH ARMY IN INDIA: ITS PRESERVATION BY AN APPROPRIATE CLOTHING, HOUSING, LOCATING, RECREATIVE EMPLOYMENT, AND HOPEFUL ENCOURAGEMENT OF THE TROOPS. with AN APPENDIX ON INDIA : THE CLIMATE OP ITS HILLS ; THE DEVELOPMENT OF ITS RESODRCBS, INDUSTRY, AND ARTS ; THE ADMINISTRATION OF JUSTICE ; THE BLACK ACT ; THE PROGRESS OF CHRISTIANITY ; THE TRAFFIC IN OPIUM ; THE VALUE OF INDIA ; PERMANENT CAUSES OF DISAFFECTION, AND OF THE RECENT REBELLION ; THE TRADITIONARY POLICY; MISGOVERNMENT BY NATIVE RULERS ; ANNEXATIONS OF THEIR TERRITORY, ETC.

THE BRITISH ARMY IN INDIA: ITS PRESERVATION BY AN APPROPRIATE CLOTHING, HOUSING, LOCATING, RECREATIVE EMPLOYMENT, AND HOPEFUL ENCOURAGEMENT OF THE TROOPS. with AN APPENDIX ON INDIA : THE CLIMATE OP ITS HILLS ; THE DEVELOPMENT OF ITS RESODRCBS, INDUSTRY, AND ARTS ; THE ADMINISTRATION OF JUSTICE ; THE BLACK ACT ; THE PROGRESS OF CHRISTIANITY ; THE TRAFFIC IN OPIUM ; THE VALUE OF INDIA ; PERMANENT CAUSES OF DISAFFECTION, AND OF THE RECENT REBELLION ; THE TRADITIONARY POLICY; MISGOVERNMENT BY NATIVE RULERS ; ANNEXATIONS OF THEIR TERRITORY, ETC.

1858 · PDF

Idries Shah 27 Books Collection : A Perfumed Scorpion, A Veiled Gazelle, Caravan of Dreams, Darkest England, Destination Mecca, Evenings with Idries Shah, Knowing How to Know, Learning How to Learn, Letters and Lectures of Idries Shah, Neglected aspects of Sufi study, Observations, Oriental Magic, Reflections, Seeker after Truth, Special Illumination, Special Problems in the study of Sufi ideas, Sufi thought and action, Tales of the Dervishes, The Dermis Probe, The Elephant in the Dark, The Englishman Handbook, Idries Shah Antology, The Magic Monastery, The natives are restless, wisdom of the Idiots PDF.

Idries Shah 27 Books Collection : A Perfumed Scorpion, A Veiled Gazelle, Caravan of Dreams, Darkest England, Destination Mecca, Evenings with Idries Shah, Knowing How to Know, Learning How to Learn, Letters and Lectures of Idries Shah, Neglected aspects of Sufi study, Observations, Oriental Magic, Reflections, Seeker after Truth, Special Illumination, Special Problems in the study of Sufi ideas, Sufi thought and action, Tales of the Dervishes, The Dermis Probe, The Elephant in the Dark, The Englishman Handbook, Idries Shah Antology, The Magic Monastery, The natives are restless, wisdom of the Idiots PDF.

2022 · PDF