# SRE Book
## 概要
"Site Reliability Engineering: How Google Runs Production Systems"(O'Reilly, 2016)は、[[Google]] が社内で培った SRE(Site Reliability Engineering)のプラクティスを体系的に公開した初の包括的文献である。[[Ben Treynor Sloss]] が創設した SRE というディシプリンを、序論・原則・実践・管理・結論の 5 部構成・全 34 章で詳述する。
## 書誌情報
- 出版社: O'Reilly Media
- 出版日: 2016-04-16
- 編者: [[Betsy Beyer]]、Chris Jones、Jennifer Petoff、[[Niall Murphy]]
- 構成: 全 34 章、5 部構成(Introduction / Principles / Practices / Management / Conclusions)
- ISBN: 978-1-491-92912-4
- URL: https://sre.google/sre-book/table-of-contents/
## 構成と主要テーマ
### Part I: Introduction(第 1〜2 章)
SRE の定義と Google における位置づけ。[[Ben Treynor Sloss]] が SRE を「ソフトウェアエンジニアに運用の設計を任せたときに生まれるもの」と定義する。
→ [[@2016__OReilly__SRE Book - Chapter 1 Introduction]]
→ [[@2016__OReilly__SRE Book - Chapter 2 The Production Environment at Google, from the Viewpoint of an SRE]] — Google 本番環境の用語法([[Borg]]・[[Chubby]]・[[Bigtable]]・RPC)を導入する用語解説の章
### Part II: Principles(第 3〜9 章)
エラーバジェット、[[サービスレベル目標]](SLO)、トイル削減、モニタリング、自動化、リリースエンジニアリング、単純性という SRE の基本原則を定義する。
→ [[@2016__OReilly__SRE Book - Chapter 3 Embracing Risk]]
→ [[@2016__OReilly__SRE Book - Chapter 4 Service Level Objectives]]
→ [[@2016__OReilly__SRE Book - Chapter 5 Eliminating Toil]]
→ [[@2016__OReilly__SRE Book - Chapter 6 Monitoring Distributed Systems]]
→ [[@2016__OReilly__SRE Book - Chapter 7 Automation at Google]]
→ [[@2016__OReilly__SRE Book - Chapter 8 Release Engineering]] — リリースエンジニアリング職能の 4 哲学と Rapid/Blaze/MPM による自動リリース基盤
→ [[@2016__OReilly__SRE Book - Chapter 9 Simplicity]] — 探索的コーディングと本質的/偶発的複雑性の区別に基づき、コード削除・最小 API・モジュール性で単純性を保つ
### Part III: Practices(第 10〜27 章)
サービス信頼性ヒエラルキーに沿った具体的プラクティスを詳述する。インシデント対応、ポストモーテム、テスト、キャパシティプランニングなどを含む。
→ [[@2016__OReilly__SRE Book - Part III Practices]]
→ [[@2016__OReilly__SRE Book - Chapter 10 Practical Alerting from Time-Series Data]] — Borgmon の設計と Prometheus への系譜、宣言型ルール評価、時系列ラベルモデル
→ [[@2016__OReilly__SRE Book - Chapter 11 Being On-Call]] — オンコールの量的・質的均衡、フォロー・ザ・サン、認知モード管理
→ [[@2016__OReilly__SRE Book - Chapter 12 Effective Troubleshooting]] — 仮説演繹法、分割統治、トリアージと安定化の優先
→ [[@2016__OReilly__SRE Book - Chapter 13 Emergency Response]] — テスト誘発型障害 vs 訓練なし障害、人間の判断力とロールバック
→ [[@2016__OReilly__SRE Book - Chapter 14 Managing Incidents]] — ICS に基づく 4 役割、フリーランシングの害、非管理型インシデントの悪化
→ [[@2016__OReilly__SRE Book - Chapter 15 Postmortem Culture - Learning from Failure]] — ブレームレス文化、経営層参加、アクションアイテムの追跡
→ [[@2016__OReilly__SRE Book - Chapter 16 Tracking Outages]] — Outalator、パッシブ集約、タグベースのメタデータ管理
→ [[@2016__OReilly__SRE Book - Chapter 17 Testing for Reliability]] — テストと信頼性の定量関係、カナリアテスト、障害の次数(U)
→ [[@2016__OReilly__SRE Book - Chapter 18 Software Engineering in SRE]] — Auxon、意図ベースのキャパシティプランニング、混合整数計画法
→ [[@2016__OReilly__SRE Book - Chapter 19 Load Balancing at the Frontend]] — フロントエンド負荷分散を DNS 層と VIP 層の 2 段構えで捉え、DNS 単独の限界(TTL 非遵守・再帰リゾルバによるクライアント位置の隠蔽・EDNS0)を示す
→ [[@2016__OReilly__SRE Book - Chapter 20 Load Balancing in the Datacenter]] — バックエンド選定を健全性判定([[レイムダック状態]])・サブセット化・負荷分散ポリシー(Weighted Round Robin)の 3 層に分解する
→ [[@2016__OReilly__SRE Book - Chapter 21 Handling Overload]] — 過負荷を QPS でなく CPU 秒で測り、クライアント側の適応スロットリングと criticality による選択的棄却を設計する
→ [[@2016__OReilly__SRE Book - Chapter 22 Addressing Cascading Failures]] — [[カスケード障害]]を正のフィードバックループとして定式化し、負荷を戻すだけでは回復しない非対称性と脱出手順を示す
→ [[@2016__OReilly__SRE Book - Chapter 23 Managing Critical State - Distributed Consensus for Reliability]] — ハートビートによるアドホックな合意の危険性と、[[分散コンセンサス]]の運用面(レプリカ配置・クォーラム構成・レイテンシ・監視)
→ [[@2016__OReilly__SRE Book - Chapter 24 Distributed Periodic Scheduling with Cron]] — べき等でない周期ジョブを分散環境で確実に一度だけ実行する問題。Paxos ログ・外部副作用の耐久化・thundering herd 回避
→ [[@2016__OReilly__SRE Book - Chapter 25 Data Processing Pipelines]] — [[周期パイプライン]]の構造的脆さ(モアレ負荷・ストラグラー)と [[Google Workflow]] による継続処理への移行
→ [[@2016__OReilly__SRE Book - Chapter 26 Data Integrity - What You Read Is What You Wrote]] — [[データ完全性]]は可用性と独立の目標であり、目的はバックアップでなく復旧である。多層防御の 3 段(ソフト削除 → バックアップと復旧 → 早期検知)
→ [[@2016__OReilly__SRE Book - Chapter 27 Reliable Product Launches at Scale]] — Launch Coordination Engineering 職能と[[ローンチチェックリスト]]。チェックリストは過去の障害から抽出された組織の記憶である
### Part IV: Management(第 28〜32 章)
SRE チームの採用・育成・組織運営・他チームとの関係構築。
→ [[@2016__OReilly__SRE Book - Chapter 28 Accelerating SRE On-Call]] — Shadow→On-Call→Project Owner の段階的オンボーディング、逆ハンドオフ、DiRT 演習
→ [[@2016__OReilly__SRE Book - Chapter 29 Dealing with Interrupts]] — 時間の二極化、コンテキストスイッチコスト、フロー状態
→ [[@2016__OReilly__SRE Book - Chapter 30 Embedding an SRE to Recover from Operational Overload]] — 学習→文脈共有→変革推進の 3 フェーズ、SLO が最重要のてこ
→ [[@2016__OReilly__SRE Book - Chapter 31 Communication and Collaboration in SRE]] — プロダクションミーティング、ハンドオフ手法、チーム構成と連携
→ [[@2016__OReilly__SRE Book - Chapter 32 The Evolving SRE Engagement Model]] — PRR→早期関与→フレームワーク、SRE プラットフォームチーム
### Part V: Conclusions(第 33〜34 章)
他産業からの教訓と結論。
→ [[@2016__OReilly__SRE Book - Chapter 33 Lessons Learned from Other Industries]] — 航空 CHIRP、医療、製造 CAPA、正常化された逸脱
→ [[@2016__OReilly__SRE Book - Chapter 34 Conclusion]]
## 刊行後の更新資料(SRE Book Updates, by Topic)
Google SRE は本書の各章について、刊行後に公開した論文・SREcon 発表・Cloud ブログ・ワークショップを [SRE Book Updates, by Topic](https://sre.google/resources/book-update/) に章別で整理している。2016 年の本文が古びた箇所を、後続資料が補っている形である。以下は各章の代表的な後続資料で、全 200 件超の一覧は原本 `.raw/books/sre-book/book-update.txt` にある。
### Part I〜II(第 2〜9 章)
- **第 2 章 本番環境**: Google's Production Environment Tech Talk(動画)
- **第 3 章 リスクの受容**: Know thy enemy: How to prioritize and communicate risks(CRE Life Lessons)
- **第 4 章 SLO**: 本書中で最も後続資料が多い章(20 件)。[[SRE Workbook]] 第 2・3・5 章、*The Calculus of Service Availability*、Cloud Architecture Center の Adopting/Defining SLOs、ワークショップ *The Art of SLOs*、書籍 *Enterprise Roadmap to SRE*
- **第 5 章 トイルの除去**: [[SRE Workbook]] 第 6 章、論文 *Invent More, Toil Less*、SREcon19 *Pragmatic Automation*
- **第 6 章 モニタリング**: *Monarch, Google's Planet-Scale Streaming Monitoring Infrastructure*(動画)、[[@2017__SREcon17 Americas__A Practical Guide to Monitoring and Alerting with Time Series at Scale]]、*The Many Ways Your Monitoring Is Lying to You*(SREcon16 EMEA)
- **第 7 章 自動化の進化**: *Prodspec and Annealing: Intent-Based Actuation for Google Production*(;login:)
- **第 8 章 リリースエンジニアリング**: [[SRE Workbook]] 第 16 章、*Making "Push on Green" a Reality*、[[@2018__acmqueue__Canary Analysis Service]]、*Canarying Well*(SREcon18 EMEA)
- **第 9 章 単純性**: [[SRE Workbook]] 第 7 章、*Relieving Technical Debt through Short Projects*(SREcon16 EMEA)
### Part III(第 10〜27 章)
- **第 10 章 実践的アラート**: [[SRE Workbook]] 第 5 章、*Reduce Toil Through Better Alerting*、[[@2018__SREcon18 Asia__A Theory and Practice of Alerting with Service Level Objectives]]、*The Structure and Interpretation of Graphs*
- **第 11 章 オンコール**: [[SRE Workbook]] 第 8 章、*Being an On-Call Engineer: A Google SRE Perspective*(;login:)、*Cognitive Bias and On-Call*(SREcon17 EMEA)
- **第 12 章 効果的なトラブルシューティング**: *Resolving Outages Faster with Better Debugging Strategies*(SREcon18 Americas)、*Traps and Cookies*
- **第 13 章 緊急対応**: [[SRE Workbook]] 第 9 章、*Debugging Incidents in Google's Distributed Systems*(ACM Queue)、*Zero Touch Prod*(SREcon19 EMEA)
- **第 14 章 インシデント管理**: *Generic mitigations: A philosophy of duct-tape outage resolution*、[[@2021__OReilly__Incident Metrics in SRE]]、*Managing Misfortune for Best Results*
- **第 15 章 ポストモーテム文化**: [[SRE Workbook]] 第 10 章・付録 C、*Postmortem Action Items: Plan the Work and Work the Plan*、共有ポストモーテムに関する CRE Life Lessons 群
- **第 16 章 障害の追跡**: *Finding the Order in Chaos*(SREcon16)
- **第 17 章 信頼性のためのテスト**: ダークローンチに関する CRE Life Lessons 2 件
- **第 18 章 SRE におけるソフトウェア工学**: *How to avoid a self-inflicted DDoS Attack*
- **第 19 章 フロントエンドの負荷分散**: [[SRE Workbook]] 第 11 章、*Maglev: A Fast and Reliable Software Network Load Balancer*、*Keeping the Balance: Internet-Scale Loadbalancing Demystified*(SREcon19)、*Anycast is Not Load Balancing*(SREcon17 EMEA)、*Doorman: Global Distributed Client Side Rate Limiting*(SREcon16)
- **第 20 章 データセンター内の負荷分散**: [[SRE Workbook]] 第 11 章、*Help Protect Your Data Centers with Safety Constraints*(SREcon18 Americas)
- **第 21 章 過負荷の処理**: [[SRE Workbook]] 第 11 章、*Using load shedding to survive a success disaster*、*Load-Shedding: Overview of Different Methodologies*(SREcon17 EMEA)
- **第 22 章 カスケード障害への対処**: [[SRE Workbook]] 第 12 章(Non-Abstract Large System Design)
- **第 23 章 クリティカルな状態の管理**: *Implementing Distributed Consensus*(SREcon20 Americas)、*Distributed Consensus Algorithms*(SREcon17 Asia、Laura Nolan)、ワークショップ *SRE Classroom: How to Design a Distributed System in 3 Hours* および *Distributed PubSub*
- **第 24 章 分散 cron**: *Reliable Cron Across the Planet*(ACM Queue。本章の元論文)
- **第 25 章 データ処理パイプライン**: [[SRE Workbook]] 第 13 章、*Reliable Data Processing with Minimal Toil*(SREcon21)、*Horizontal Data Freshness Monitoring in Complex Pipelines*(SREcon21)、*Distributed Log-Processing Design Workshop*
- **第 26 章 データ完全性**: *Tradeoffs in Resiliency: Managing the Burden of Data Recoverability*(SREcon18 EMEA)
- **第 27 章 大規模での信頼できるローンチ**: *Reliable Launches at Scale*(SREcon17 Asia、Sebastian Kirsch)
### Part IV〜V(第 28〜33 章)
- **第 28 章 オンコールへの加速**: SRE 教育プログラムに関する SREcon 発表群(*Deploying SRE Training Best Practices to Production* ほか)、*Interviewing for Systems Design Skills*
- **第 29 章 割り込みへの対処**: 論文 *Interrupt Reduction Projects*
- **第 30 章 運用過負荷からの回復**: [[SRE Workbook]] 第 17 章
- **第 31 章 コミュニケーションと協業**: [[SRE Workbook]] 付録 A(SLO 文書例)・付録 B(エラーバジェット方針例)、エスカレーション方針に関する CRE Life Lessons、ワークショップ *Developing a Google SRE Culture*
- **第 32 章 SRE エンゲージメントモデルの進化**: [[SRE Workbook]] 第 18 章、*Bigtable: A Journey from Binary to Service*(SREcon19 EMEA)
- **第 33 章 他産業からの教訓**: *A Political Scientist's View on Site Reliability*(SREcon21)、*SRE for Good*(SREcon18 EMEA)
> [!note] 第 1 章・第 34 章・付録には後続資料の登録がない。
## 影響と位置づけ
本書の出版は SRE を Google 固有の実践から業界標準のディシプリンへと転換した画期的な出来事である。エラーバジェット、ブレームレスポストモーテム、トイルの定量化、4 つのゴールデンシグナルなどの概念は、本書を通じて広く普及した。
## 関連書籍
- **The Site Reliability Workbook**(O'Reilly, 2018): 実践的なハウツーを補完するコンパニオン書籍。[[Betsy Beyer]]、[[Niall Murphy]] らが編者を務める。
- **Building Secure & Reliable Systems**(O'Reilly, 2020): セキュリティと信頼性の統合を論じた関連書籍。
## SRE Workbook との関係
[[SRE Workbook]] は本書の原則を導入手順へ落とす続編である。SRE Book が SLO・エラーバジェット・トイル・インシデント管理・ポストモーテム文化の語彙を定義したのに対し、Workbook は SLI 仕様/実装、SLO 文書、エラーバジェット方針、複数ウィンドウ複数バーン率アラート、オンコール負荷管理、ポストモーテムテンプレートへ具体化する。
## 関連
- [[@2007__LISA__On Designing and Deploying Internet-Scale Services]]: 本書の思想的先駆者の一つ
- [[@1983__Automatica__Ironies of Automation]]: 自動化の章で参照される古典
- [[Google]]: 本書の母体組織
## 出典
- Betsy Beyer, Chris Jones, Jennifer Petoff, Niall Murphy (eds.), *Site Reliability Engineering: How Google Runs Production Systems*, O'Reilly, 2016