
Site Reliability Engineering
Description
Book Introduction
A book full of inspiration to help you grow to the next level!
Experience Google's principle that code that actually works is the most important!
The lifespan of a software system is usually determined by the period during which it is actually used, not by its design or implementation phase.
So why have software engineers considered the process of designing and implementing large-scale computing systems to be of the utmost importance?
In this book, core members of Google's Site Reliability Engineering team share, through essays and editorials, how and why they implement, deploy, monitor, and maintain the world's largest software system by focusing on the entire software lifecycle.
This will allow you to apply the principles and practices that enabled Google engineers to build more scalable, reliable, and effective systems to your own organization.
Main contents of this book
*Introduction: Introduces what site reliability engineering is and how it differs from traditional IT practices.
*Principles: Introduces patterns, behaviors, and various issues that affect the work of site reliability engineers.
*Case Study: Learn the theory and case studies for building and operating large-scale distributed computing systems, which is part of SRE's daily work.
*Management: Explore ways to apply Google's recommended new hire training, communication, and meeting management practices to your own organization.
Experience Google's principle that code that actually works is the most important!
The lifespan of a software system is usually determined by the period during which it is actually used, not by its design or implementation phase.
So why have software engineers considered the process of designing and implementing large-scale computing systems to be of the utmost importance?
In this book, core members of Google's Site Reliability Engineering team share, through essays and editorials, how and why they implement, deploy, monitor, and maintain the world's largest software system by focusing on the entire software lifecycle.
This will allow you to apply the principles and practices that enabled Google engineers to build more scalable, reliable, and effective systems to your own organization.
Main contents of this book
*Introduction: Introduces what site reliability engineering is and how it differs from traditional IT practices.
*Principles: Introduces patterns, behaviors, and various issues that affect the work of site reliability engineers.
*Case Study: Learn the theory and case studies for building and operating large-scale distributed computing systems, which is part of SRE's daily work.
*Management: Explore ways to apply Google's recommended new hire training, communication, and meeting management practices to your own organization.
- You can preview some of the book's contents.
Preview
index
PART I INTRODUCTION
CHAPTER 01 Introduction _ 3
How to Use System Administrators for Service Management _ 3
Google's Solution to Service Management: Site Reliability Engineering _ 5
SRE's Creed _ 8
In conclusion _ 14
CHAPTER 02 Google's Production Environment from an SRE Perspective 15
Hardware _ 15
System software that 'tunes' the hardware _ 17
Other System Software _ 21
Software Infrastructure _ 22
Development Environment _ 23
Shakespeare: Example Service _ 24
PART II PRINCIPLES AND RULES
CHAPTER 03 Accepting Risks _ 30
Managing Risk Factors _ 31
Measuring Service Risk _ 32
Service Risk Acceptance _ 34
Using the Error Budget _ 40
CHAPTER 04 Service Level Objectives _ 44
Service Level Related Terminology _ 45
Indicator Settings _ 48
Goal Setting Exercise _ 51
Practice on the Agreement _ 56
CHAPTER 05 No More Digging! _ 57
The Definition of Digging _ 58
Reasons Why Less Digging is Good _ 60
What kind of work falls under engineering? _ 61
Is digging always a bad thing? _ 62
Conclusion _ 63
CHAPTER 06 Distributed System Monitoring _ 64
Definition _ 64
Why Monitor? _ 66
Setting Appropriate Expectations for Monitoring _ 67
Symptoms and Causes _ 68
Black Box and White Box _ 69
Four Crucial Indicators _ 70
Considerations for the last request (or execution and performance) _ 72
Choosing the Right Measurement Method _ 72
Not simpler, but as simple as possible _ 73
Combining the principles we've discussed so far _ 74
Long-term monitoring _ 76
Conclusion _ 78
CHAPTER 07 Google's Advanced Automation _ 80
The Value of Automation _ 81
The Value of Google SRE _ 83
Automation Case Study _ 85
Do Yourself a favor: Automate Everything! _ 88
A Move of God: Automating Cluster Turn-Ups _ 91
Vogue: The Birth of a Warehouse-Sized Computer _ 98
Reliability is a fundamental function _ 101
Recommendations _ 101
CHAPTER 08 Release Engineering _ 103
The Role of a Release Engineer _ 104
The Philosophy of Release Engineering _ 104
Continuous Build and Deployment _ 106
Configuration Management Techniques _ 111
Conclusion _ 112
CHAPTER 09 Conciseness _ 114
System stability vs.
Swiftness _ 115
The Virtue of Boredom _ 115
My code will never give up! _ 116
Indicator of 'Negative Impact Code' _ 116
Minimal API _ 117
Modularization _ 117
Simplifying Releases _ 118
Concise Conclusion _ 119
PART III CASE
CHAPTER 10 Practical Reminders for Time Series Data _ 127
The Birth of Borgmon _ 128
Manipulating the Application _ 130
Collection of exported data _ 131
Storage for Time Series Data _ 132
Evaluation of Rules _ 135
Notification _ 140
Sharding in Monitoring Topologies _ 141
Black Box Monitoring _ 142
Maintaining Settings _ 143
Over the past 10 years… _ 145
CHAPTER 11 Emergency Standby _ 146
Introduction _ 146
The Life of an Emergency Response Engineer _ 147
Balancing Emergency Standby Work _ 148
Considering Safety _ 150
Escape from Inappropriate Operating Loads _ 153
Conclusion _ 155
CHAPTER 12 Effective Failover _ 156
Theory _ 157
Let's get into the real world _ 159
The Magic of Negative Results _ 168
Case Study _ 171
Making Handling Disabilities a Little Easier _ 175
Conclusion _ 176
CHAPTER 13 EMERGENCY RESPONSE _ 177
What should I do if something goes wrong with my system? _ 178
Test-induced failures _ 178
Disability due to change _ 180
Procedural Obstacles _ 183
All problems solved _ 186
Learn from the past.
And don't repeat _ 186
Conclusion _ 188
CHAPTER 14 MANAGING DISABILITIES _ 189
Inadequate Disability Management _ 190
A Detailed Analysis of Inadequate Handling of Disabilities _ 191
Fundamentals of Disability Management Procedures _ 191
Properly Managed Failover _ 194
When to declare a disability? _ 195
Summary _ 196
CHAPTER 15 Postmortem Culture: Learning from Failure _ 197
Google's Postmortem Philosophy _ 198
Collaboration and Knowledge Sharing _ 200
Introducing Postmortem Culture _ 201
Conclusion and Continuous Improvement _ 204
CHAPTER 16 TRACING SYSTEM OUTAGES _ 205
Escalator _ 206
Outerator _ 206
CHAPTER 17 Testing for Reliability _ 212
Types of Software Testing _ 214
Configuring the Test and Build Environment _ 221
Testing in Large-Scale Environments _ 223
Conclusion _ 237
CHAPTER 18 Software Engineering in SRE Organizations _ 238
Why Software Engineering Capabilities Matter in SRE Organizations _ 239
Auxon Case Study: Project Background and Where the Problems Arrived _ 240
Intent-Based Capacity Planning _ 244
How to Foster Software Engineering in Your SRE Organization _ 254
Conclusion _ 259
CHAPTER 19 Front-End Load Balancing _ 260
Not everything can be solved by force alone _ 260
Load Balancing Using DNS _ 262
Load Balancing Using Virtual IP Addresses _ 265
CHAPTER 20 Load Balancing in Data Centers _ 268
Ideal Case _ 269
Distinguishing Bad Tasks: Flow Control and Lame Ducks _ 271
Limiting the Connection Pool Using Subsets _ 273
Load Balancing Policy _ 280
CHAPTER 21 Handling Overload _ 287
The Pitfalls of 'Query Per Second' _ 288
Per-user limits _ 289
Client-side usage limits _ 290
Importance _ 292
Signals of Utilization _ 294
Handling Overload Errors _ 295
Load on connection _ 299
Conclusion _ 300
CHAPTER 22 Handling Continuous Failures _ 302
Causes of Continuous Disability and Countermeasures _ 303
Preventing Server Overload _ 309
Slow Start and Cold Caching _ 320
Causes of continuous disability _ 323
Testing for Continuous Failures _ 325
Immediate Response to Continuous Failures _ 328
In conclusion _ 331
CHAPTER 23 Managing Critical States: Agreeing on Distributed Reliability _ 332
Why Consensus Is Necessary: Failures in Collaboration Across Distributed Systems _ 335
How Consensus on Distributed Blockchain Works _ 337
System Architecture Patterns for Distributed Consensus _ 339
Performance of Distributed Consensus _ 345
Deploying a Distributed Consensus-Based System _ 354
Distributed Consensus System Monitoring _ 364
Conclusion _ 365
CHAPTER 24 Distributed Periodic Scheduling with Cron _ 366
Cron_367
Cron Jobs and Idempotence _ 368
Cron in Large Systems _ 369
Cron Service Implemented by Google _ 371
Summary _ 379
CHAPTER 25 DATA PROCESSING PIPELINE _ 380
The Origins of the Pipeline Design Pattern _ 380
The Basic Effects of Big Data Using a Simple Pipeline Pattern _ 381
Challenges of the Regular Pipeline Pattern _ 381
Problems arising from unbalanced distribution of work _ 382
Drawbacks of Regular Pipelines in Distributed Environments _ 383
Introducing Google Workflow _ 387
Workflow Execution Steps _ 389
Ensuring Business Sustainability _ 391
Summary _ 392
CHAPTER 26 DATA INTEGRITY: I MUST BE READY TO READ WHAT I WRITE _ 394
Critical Conditions for Data Integrity _ 395
Google SRE's Goals for Maintaining Data Integrity and Availability _ 401
How Google Solves Data Integrity Problems _ 406
Case Study _ 419
General Principles of SRE Related to Data Integrity _ 427
Conclusion _ 429
CHAPTER 27: Delivering Reliable Products in High-Volume Environments _ 430
Release Coordination Engineering _ 432
Establishing a Launch Procedure _ 434
Developing a Launch Checklist _ 438
Techniques for a Stable Release _ 443
LCE's Development Capabilities _ 448
Conclusion _ 452
PART IV MANAGEMENT
CHAPTER 28: Fostering SRE Growth Beyond Emergency Standby _ 456
Hired a new SRE.
What should I do now? _ 456
First Learning Experience: A Case Study of Structure to Prevent Confusion _ 459
Raising Star Reverse Engineers and Improvisational Thinkers _ 463
Five Principles for Aspiring Emergency Response Engineers _ 467
Beyond Emergency Standby: Rites of Passage and Practices of Continuous Learning _ 473
In conclusion _ 474
CHAPTER 29 Dealing with Distractions _ 475
Managing Operational Workloads _ 476
Factors for Determining How to Manage Distractions _ 477
Imperfect Machine _ 478
CHAPTER 30: Freeing Yourself from the Burden of Operational Tasks with SRE _ 486
Step 1: Learn about the service and understand the context _ 487
Step 2: Sharing Context _ 489
Step 3: Leading the Change _ 491
Conclusion _ 494
CHAPTER 31 Communication and Collaboration in SRE _ 495
Communication: Operational Environment Meeting _ 497
Collaborating with SREs _ 501
A Case Study on Collaboration in SRE: Viceroy _ 503
Collaboration with Organizations Outside of SRE _ 509
Case Study: DFP's Migration to F1 _ 510
Conclusion _ 512
CHAPTER 32 Improving the SRE Participation Model _ 513
Introducing SRE: What It Is, How It Works, and Why It Works _ 513
PRR Model _ 514
SRE Introduction Model _ 515
Reviewing the Operational Environment Readiness: A Simple PRR Model _ 517
Evolution of a Simple PRR Model: Early Engagement _ 521
Improving Service Development: Frameworks and SRE Platforms _ 524
Conclusion _ 529
PART V CONCLUSION
CHAPTER 33 Lessons from Other Industries _ 533
Introduction to Industry Experts _ 534
Preparation and Disaster Testing _ 536
Postmortem Culture _ 540
Eliminate repetitive tasks and operational overhead through automation _ 542
Structured and Rational Decision Making _ 544
Conclusion _ 546
CHAPTER 34 Conclusion _ 547
APPENDIX.
supplement
APPENDIX A Availability Table _ 551
APPENDIX B Collection of Recommended Practices for Operational Services _ 553
APPENDIX C Example of a Failure Status Document _ 559
APPENDIX D Postmortem Examples _ 561
APPENDIX E Launch Coordination Checklist _ 566
APPENDIX F Example of a Product Meeting _ 569
References _ 573
Search _ 584
CHAPTER 01 Introduction _ 3
How to Use System Administrators for Service Management _ 3
Google's Solution to Service Management: Site Reliability Engineering _ 5
SRE's Creed _ 8
In conclusion _ 14
CHAPTER 02 Google's Production Environment from an SRE Perspective 15
Hardware _ 15
System software that 'tunes' the hardware _ 17
Other System Software _ 21
Software Infrastructure _ 22
Development Environment _ 23
Shakespeare: Example Service _ 24
PART II PRINCIPLES AND RULES
CHAPTER 03 Accepting Risks _ 30
Managing Risk Factors _ 31
Measuring Service Risk _ 32
Service Risk Acceptance _ 34
Using the Error Budget _ 40
CHAPTER 04 Service Level Objectives _ 44
Service Level Related Terminology _ 45
Indicator Settings _ 48
Goal Setting Exercise _ 51
Practice on the Agreement _ 56
CHAPTER 05 No More Digging! _ 57
The Definition of Digging _ 58
Reasons Why Less Digging is Good _ 60
What kind of work falls under engineering? _ 61
Is digging always a bad thing? _ 62
Conclusion _ 63
CHAPTER 06 Distributed System Monitoring _ 64
Definition _ 64
Why Monitor? _ 66
Setting Appropriate Expectations for Monitoring _ 67
Symptoms and Causes _ 68
Black Box and White Box _ 69
Four Crucial Indicators _ 70
Considerations for the last request (or execution and performance) _ 72
Choosing the Right Measurement Method _ 72
Not simpler, but as simple as possible _ 73
Combining the principles we've discussed so far _ 74
Long-term monitoring _ 76
Conclusion _ 78
CHAPTER 07 Google's Advanced Automation _ 80
The Value of Automation _ 81
The Value of Google SRE _ 83
Automation Case Study _ 85
Do Yourself a favor: Automate Everything! _ 88
A Move of God: Automating Cluster Turn-Ups _ 91
Vogue: The Birth of a Warehouse-Sized Computer _ 98
Reliability is a fundamental function _ 101
Recommendations _ 101
CHAPTER 08 Release Engineering _ 103
The Role of a Release Engineer _ 104
The Philosophy of Release Engineering _ 104
Continuous Build and Deployment _ 106
Configuration Management Techniques _ 111
Conclusion _ 112
CHAPTER 09 Conciseness _ 114
System stability vs.
Swiftness _ 115
The Virtue of Boredom _ 115
My code will never give up! _ 116
Indicator of 'Negative Impact Code' _ 116
Minimal API _ 117
Modularization _ 117
Simplifying Releases _ 118
Concise Conclusion _ 119
PART III CASE
CHAPTER 10 Practical Reminders for Time Series Data _ 127
The Birth of Borgmon _ 128
Manipulating the Application _ 130
Collection of exported data _ 131
Storage for Time Series Data _ 132
Evaluation of Rules _ 135
Notification _ 140
Sharding in Monitoring Topologies _ 141
Black Box Monitoring _ 142
Maintaining Settings _ 143
Over the past 10 years… _ 145
CHAPTER 11 Emergency Standby _ 146
Introduction _ 146
The Life of an Emergency Response Engineer _ 147
Balancing Emergency Standby Work _ 148
Considering Safety _ 150
Escape from Inappropriate Operating Loads _ 153
Conclusion _ 155
CHAPTER 12 Effective Failover _ 156
Theory _ 157
Let's get into the real world _ 159
The Magic of Negative Results _ 168
Case Study _ 171
Making Handling Disabilities a Little Easier _ 175
Conclusion _ 176
CHAPTER 13 EMERGENCY RESPONSE _ 177
What should I do if something goes wrong with my system? _ 178
Test-induced failures _ 178
Disability due to change _ 180
Procedural Obstacles _ 183
All problems solved _ 186
Learn from the past.
And don't repeat _ 186
Conclusion _ 188
CHAPTER 14 MANAGING DISABILITIES _ 189
Inadequate Disability Management _ 190
A Detailed Analysis of Inadequate Handling of Disabilities _ 191
Fundamentals of Disability Management Procedures _ 191
Properly Managed Failover _ 194
When to declare a disability? _ 195
Summary _ 196
CHAPTER 15 Postmortem Culture: Learning from Failure _ 197
Google's Postmortem Philosophy _ 198
Collaboration and Knowledge Sharing _ 200
Introducing Postmortem Culture _ 201
Conclusion and Continuous Improvement _ 204
CHAPTER 16 TRACING SYSTEM OUTAGES _ 205
Escalator _ 206
Outerator _ 206
CHAPTER 17 Testing for Reliability _ 212
Types of Software Testing _ 214
Configuring the Test and Build Environment _ 221
Testing in Large-Scale Environments _ 223
Conclusion _ 237
CHAPTER 18 Software Engineering in SRE Organizations _ 238
Why Software Engineering Capabilities Matter in SRE Organizations _ 239
Auxon Case Study: Project Background and Where the Problems Arrived _ 240
Intent-Based Capacity Planning _ 244
How to Foster Software Engineering in Your SRE Organization _ 254
Conclusion _ 259
CHAPTER 19 Front-End Load Balancing _ 260
Not everything can be solved by force alone _ 260
Load Balancing Using DNS _ 262
Load Balancing Using Virtual IP Addresses _ 265
CHAPTER 20 Load Balancing in Data Centers _ 268
Ideal Case _ 269
Distinguishing Bad Tasks: Flow Control and Lame Ducks _ 271
Limiting the Connection Pool Using Subsets _ 273
Load Balancing Policy _ 280
CHAPTER 21 Handling Overload _ 287
The Pitfalls of 'Query Per Second' _ 288
Per-user limits _ 289
Client-side usage limits _ 290
Importance _ 292
Signals of Utilization _ 294
Handling Overload Errors _ 295
Load on connection _ 299
Conclusion _ 300
CHAPTER 22 Handling Continuous Failures _ 302
Causes of Continuous Disability and Countermeasures _ 303
Preventing Server Overload _ 309
Slow Start and Cold Caching _ 320
Causes of continuous disability _ 323
Testing for Continuous Failures _ 325
Immediate Response to Continuous Failures _ 328
In conclusion _ 331
CHAPTER 23 Managing Critical States: Agreeing on Distributed Reliability _ 332
Why Consensus Is Necessary: Failures in Collaboration Across Distributed Systems _ 335
How Consensus on Distributed Blockchain Works _ 337
System Architecture Patterns for Distributed Consensus _ 339
Performance of Distributed Consensus _ 345
Deploying a Distributed Consensus-Based System _ 354
Distributed Consensus System Monitoring _ 364
Conclusion _ 365
CHAPTER 24 Distributed Periodic Scheduling with Cron _ 366
Cron_367
Cron Jobs and Idempotence _ 368
Cron in Large Systems _ 369
Cron Service Implemented by Google _ 371
Summary _ 379
CHAPTER 25 DATA PROCESSING PIPELINE _ 380
The Origins of the Pipeline Design Pattern _ 380
The Basic Effects of Big Data Using a Simple Pipeline Pattern _ 381
Challenges of the Regular Pipeline Pattern _ 381
Problems arising from unbalanced distribution of work _ 382
Drawbacks of Regular Pipelines in Distributed Environments _ 383
Introducing Google Workflow _ 387
Workflow Execution Steps _ 389
Ensuring Business Sustainability _ 391
Summary _ 392
CHAPTER 26 DATA INTEGRITY: I MUST BE READY TO READ WHAT I WRITE _ 394
Critical Conditions for Data Integrity _ 395
Google SRE's Goals for Maintaining Data Integrity and Availability _ 401
How Google Solves Data Integrity Problems _ 406
Case Study _ 419
General Principles of SRE Related to Data Integrity _ 427
Conclusion _ 429
CHAPTER 27: Delivering Reliable Products in High-Volume Environments _ 430
Release Coordination Engineering _ 432
Establishing a Launch Procedure _ 434
Developing a Launch Checklist _ 438
Techniques for a Stable Release _ 443
LCE's Development Capabilities _ 448
Conclusion _ 452
PART IV MANAGEMENT
CHAPTER 28: Fostering SRE Growth Beyond Emergency Standby _ 456
Hired a new SRE.
What should I do now? _ 456
First Learning Experience: A Case Study of Structure to Prevent Confusion _ 459
Raising Star Reverse Engineers and Improvisational Thinkers _ 463
Five Principles for Aspiring Emergency Response Engineers _ 467
Beyond Emergency Standby: Rites of Passage and Practices of Continuous Learning _ 473
In conclusion _ 474
CHAPTER 29 Dealing with Distractions _ 475
Managing Operational Workloads _ 476
Factors for Determining How to Manage Distractions _ 477
Imperfect Machine _ 478
CHAPTER 30: Freeing Yourself from the Burden of Operational Tasks with SRE _ 486
Step 1: Learn about the service and understand the context _ 487
Step 2: Sharing Context _ 489
Step 3: Leading the Change _ 491
Conclusion _ 494
CHAPTER 31 Communication and Collaboration in SRE _ 495
Communication: Operational Environment Meeting _ 497
Collaborating with SREs _ 501
A Case Study on Collaboration in SRE: Viceroy _ 503
Collaboration with Organizations Outside of SRE _ 509
Case Study: DFP's Migration to F1 _ 510
Conclusion _ 512
CHAPTER 32 Improving the SRE Participation Model _ 513
Introducing SRE: What It Is, How It Works, and Why It Works _ 513
PRR Model _ 514
SRE Introduction Model _ 515
Reviewing the Operational Environment Readiness: A Simple PRR Model _ 517
Evolution of a Simple PRR Model: Early Engagement _ 521
Improving Service Development: Frameworks and SRE Platforms _ 524
Conclusion _ 529
PART V CONCLUSION
CHAPTER 33 Lessons from Other Industries _ 533
Introduction to Industry Experts _ 534
Preparation and Disaster Testing _ 536
Postmortem Culture _ 540
Eliminate repetitive tasks and operational overhead through automation _ 542
Structured and Rational Decision Making _ 544
Conclusion _ 546
CHAPTER 34 Conclusion _ 547
APPENDIX.
supplement
APPENDIX A Availability Table _ 551
APPENDIX B Collection of Recommended Practices for Operational Services _ 553
APPENDIX C Example of a Failure Status Document _ 559
APPENDIX D Postmortem Examples _ 561
APPENDIX E Launch Coordination Checklist _ 566
APPENDIX F Example of a Product Meeting _ 569
References _ 573
Search _ 584
Into the book
Beyond that, the belief and talent for solving complex problems through software system development were cited as essential qualities for SREs. Within the SRE team, we closely tracked the development of these two groups' work capabilities, and to date, we have found no significant differences in the performance of these two groups of engineers.
In fact, the SRE team's background differences often led to the creation of unique, high-quality systems that integrated multiple technologies.
--- p.6
Understanding how well your system is meeting expectations helps you decide whether to invest in making it faster, more available, and more resilient.
Alternatively, if the system is working well, the staff's time can be allocated to higher-priority tasks, such as addressing technical debt, adding new features, or developing other products.
--- p.56
The system has inertia.
We found that a computer system that is functioning properly tends to continue to function until an external factor, such as a change in configuration or a change in the type of service load, occurs.
So the best place to start figuring out what's going wrong is the most recent change.
--- p.165
A reasonable way to implement a subset selection algorithm is to have each client randomly shuffle the list of backends, then select the backends that are accessible and in good condition to build a subset.
The method of randomly shuffling and then selecting backends sequentially allows us to explicitly limit the number of backends to be considered, thus handling restarts and failures reliably (e.g. while maintaining a relatively small number of connections).
However, we have experienced that this strategy does not work as desired in most cases because the load is not distributed evenly.
--- p.275
Companies that are growing rapidly and have a high rate of change in their products and services may benefit from introducing a release orchestration engineering role.
Such teams are especially useful when a company plans to hire more than twice as many product developers every year or two, when a service must scale to millions of users, or when stability is more important to users despite a high rate of change.
In fact, the SRE team's background differences often led to the creation of unique, high-quality systems that integrated multiple technologies.
--- p.6
Understanding how well your system is meeting expectations helps you decide whether to invest in making it faster, more available, and more resilient.
Alternatively, if the system is working well, the staff's time can be allocated to higher-priority tasks, such as addressing technical debt, adding new features, or developing other products.
--- p.56
The system has inertia.
We found that a computer system that is functioning properly tends to continue to function until an external factor, such as a change in configuration or a change in the type of service load, occurs.
So the best place to start figuring out what's going wrong is the most recent change.
--- p.165
A reasonable way to implement a subset selection algorithm is to have each client randomly shuffle the list of backends, then select the backends that are accessible and in good condition to build a subset.
The method of randomly shuffling and then selecting backends sequentially allows us to explicitly limit the number of backends to be considered, thus handling restarts and failures reliably (e.g. while maintaining a relatively small number of connections).
However, we have experienced that this strategy does not work as desired in most cases because the load is not distributed evenly.
--- p.275
Companies that are growing rapidly and have a high rate of change in their products and services may benefit from introducing a release orchestration engineering role.
Such teams are especially useful when a company plans to hire more than twice as many product developers every year or two, when a service must scale to millions of users, or when stability is more important to users despite a high rate of change.
--- p.452
GOODS SPECIFICS
- Date of publication: January 18, 2018
- Page count, weight, size: 624 pages | 188*245*29mm
- ISBN13: 9791188621088
- ISBN10: 1188621084
You may also like
카테고리
korean
korean