# Building Secure and Reliable Systems ## 概要 Google の Security 部門と SRE 部門の実務者およそ 150 名が共同執筆した、セキュリティと信頼性を「後付けの品質」ではなくシステムライフサイクル全体に組み込むべき第一級の設計要件として扱う書籍である。既刊の [[SRE Book]](2016)・[[SRE Workbook]](2018)が信頼性を扱いながら踏み込まなかった「信頼性とセキュリティの交差点」を主題に据え、設計(Part II)・実装(Part III)・運用(Part IV)・組織と文化(Part V)という開発ライフサイクルの順で構成する。(Source: [[Building Secure and Reliable Systems]] Preface) 本書の中心的な主張は、信頼性とセキュリティがいずれもシステムの**創発特性**であって後付けできず、両者は共通点(不可視性・ライフサイクル全体にわたる関与・複雑性への対処)を多く持ちながら、**敵対者の有無**という一点で設計判断が分岐する、というものである。(Source: [[@2020__OReilly__Building Secure and Reliable Systems - Chapter 1 The Intersection of Security and Reliability]]) ## 書誌情報 - 原題: *Building Secure and Reliable Systems: Best Practices for Designing, Implementing, and Maintaining Systems* - 編者: Heather Adkins, Betsy Beyer, Paul Blankinship, Piotr Lewandowski, Ana Oprea, Adam Stubblefield - 出版社: O'Reilly Media - 発行: 2020 年 3 月 - ISBN: 978-1-492-08312-2 - 構成: 全 5 部 21 章 + Conclusion + Appendix(付録 A: 災害リスク評価マトリクス) - 公開版: https://google.github.io/building-secure-and-reliable-systems/raw/toc.html(全文が無償公開されている) - 序文: Royal Hansen(Vice President, Security Engineering)、Michael Wildpaner(Senior Director, Site Reliability Engineering) ## 構成と主要テーマ ### Part I. Introductory Material(第 1〜2 章) 信頼性とセキュリティが何を共有し、どこで分岐するかを提示し、次いで「敵対者とは誰か」を枠組み化して本書全体の前提を置く。 → [[@2020__OReilly__Building Secure and Reliable Systems - Chapter 1 The Intersection of Security and Reliability]] — 信頼性とセキュリティは共に創発特性だが、敵対者の有無という一点が両者の設計判断を分岐させる → [[@2020__OReilly__Building Secure and Reliable Systems - Chapter 2 Understanding Adversaries]] — 敵対者を動機・プロファイル・手法の 3 枠組みで捉え、内部者リスクへの設計対策を示す ### Part II. Designing Systems(第 3〜10 章) セキュリティと信頼性を最も安く実現できるのは設計段階だという立場から、既存システムへの後付け(第 3 章)から始め、要件のトレードオフ、最小権限、理解容易性、変化への追随、レジリエンス、復旧、DoS 緩和へと進む。 → [[@2020__OReilly__Building Secure and Reliable Systems - Chapter 3 Case Study - Safe Proxies]] — セーフプロキシで本番操作を仲介し、監査・多者承認・レート制限を後付けする設計手法 → [[@2020__OReilly__Building Secure and Reliable Systems - Chapter 4 Design Tradeoffs]] — 機能要件とセキュリティ・信頼性は対立せず、初期投資により両立できると説く → [[@2020__OReilly__Building Secure and Reliable Systems - Chapter 5 Design for Least Privilege]] — 最小権限を分類・小さな API・監査・多者承認等の高度な制御に分解して実装する → [[@2020__OReilly__Building Secure and Reliable Systems - Chapter 6 Design for Understandability]] — 不変条件・TCB・アイデンティティ層・型で複雑性を封じ込め、理解容易性を設計品質として扱う → [[@2020__OReilly__Building Secure and Reliable Systems - Chapter 7 Design for a Changing Landscape]] — 短期・中期・長期の変化に高信頼性を保ったまま追随する設計と、HTTPS 移行という業界規模の移行の完遂法 → [[@2020__OReilly__Building Secure and Reliable Systems - Chapter 8 Design for Resilience]] — 多層防御・劣化制御・影響範囲の分離・障害ドメイン冗長化・継続的検証でレジリエンスを設計する → [[@2020__OReilly__Building Secure and Reliable Systems - Chapter 9 Design for Recovery]] — 復旧は速度とポリシーの分離、単調な MASVN によるロールバック制御、意図した状態の把握で支える → [[@2020__OReilly__Building Secure and Reliable Systems - Chapter 10 Mitigating Denial-of-Service Attacks]] — DoS を経済的非対称と捉え直し、階層防御・自動緩和・自作の攻撃への対処を論じる ### Part III. Implementing Systems(第 11〜15 章) 設計を実装に落とす段階を扱う。買うか作るかの判断(第 11 章)から、コードを書く・テストする・デプロイする各段階の安全性、そして問題が起きたときに調べる能力までを順に論じる。 → [[@2020__OReilly__Building Secure and Reliable Systems - Chapter 11 Case Study - Designing, Implementing, and Maintaining a Publicly Trusted CA]] — 公的信頼 CA を内製する判断と、その設計・鍵保護・発行検証の運用 → [[@2020__OReilly__Building Secure and Reliable Systems - Chapter 12 Writing Code]] — 型とフレームワークで危険なコードをコンパイル不能にし、開発者の注意力に頼らず安全性を構造的に強制する → [[@2020__OReilly__Building Secure and Reliable Systems - Chapter 13 Testing Code]] — 単体・統合・動的解析・ファジング・静的解析を開発者ワークフローへ統合して初めて効果が積み上がる → [[@2020__OReilly__Building Secure and Reliable Systems - Chapter 14 Deploying Code]] — デプロイは「誰が」ではなく「何を」検証すべきであり、バイナリ来歴・検証可能ビルド・チョークポイントで敵対者の迂回を防ぐ → [[@2020__OReilly__Building Secure and Reliable Systems - Chapter 15 Investigating Systems]] — デバッグ技法とセキュリティ調査の違い、およびログの不変性・保持・アクセス制御の設計トレードオフ ### Part IV. Maintaining Systems(第 16〜18 章) インシデントを時間軸で 3 分割し、事前準備・危機の最中・事後の復旧をそれぞれ独立した章として扱う。 → [[@2020__OReilly__Building Secure and Reliable Systems - Chapter 16 Disaster Planning]] — 災害が起きる前の準備。リスク分析・IR チーム組成・重大度/優先度モデル・段階的テスト → [[@2020__OReilly__Building Secure and Reliable Systems - Chapter 17 Crisis Management]] — 危機か否かのトリアージから運用セキュリティ・並列化・引き継ぎ・士気管理まで、危機の指揮統制 → [[@2020__OReilly__Building Secure and Reliable Systems - Chapter 18 Recovery and Aftermath]] — セキュリティ復旧は攻撃者という反応する相手を前提に、追い出しの時機と技術的負債のトレードオフを設計する ### Part V. Organization and Culture(第 19〜21 章) 技術的実践は組織文化に支えられて初めて機能するという立場から、具体的なチームの事例、役割分担の一般原則、文化の設計へと進む。 → [[@2020__OReilly__Building Secure and Reliable Systems - Chapter 19 Case Study - Chrome Security Team]] — Chrome セキュリティチームは脆弱性報奨金プログラム発足とハイブリッドエンジニアリングチーム化を経て多層防御と透明性の文化を築いた → [[@2020__OReilly__Building Secure and Reliable Systems - Chapter 20 Understanding Roles and Responsibilities]] — セキュリティは全員の責任であり、専門家は専門実装とベストプラクティス整備に徹すべきだと説く → [[@2020__OReilly__Building Secure and Reliable Systems - Chapter 21 Building a Culture of Security and Reliability]] — 本書全体の技術的実践は、意図的に設計された組織文化に支えられて初めて機能する ### 結びと付録 → [[@2020__OReilly__Building Secure and Reliable Systems - Conclusion]] — セキュリティと信頼性はシステム固有の属性であり全員の責任だと結び、知識領域の横断とチームへの投資を訴える → [[@2020__OReilly__Building Secure and Reliable Systems - Appendix A Disaster Risk Assessment Matrix]] — 発生確率(6 段階)×影響度(5 段階)で災害リスクを格付けするサンプルマトリクス ## 影響と位置づけ [[SRE Book]](2016)と [[SRE Workbook]](2018)が信頼性の原則と実践を扱ったのに対し、本書は同じ Google の実務者コミュニティから「信頼性とセキュリティの交差点」を主題として書かれた。序文で Royal Hansen(Vice President, Security Engineering)は、SRE が取り組む問題空間とセキュリティの問題空間が似た力学を示すこと、SRE が「役割と責務の定義」に加えて「チームを接続する実装モデル」を作った点こそセキュリティコミュニティが次に踏むべき一歩だと述べる。Michael Wildpaner(Senior Director, Site Reliability Engineering)は、SRE を「悪い設計や悪い実装がシステムのセキュリティに影響するのを防ぐ最後の防衛線」と位置づける。(Source: [[Building Secure and Reliable Systems]] Foreword) 本書の特徴は、抽象的な原則ではなく Google の内部インフラで実際に運用されている仕組み——[[Zero Touch Prod]]・[[Google Tool Proxy]]・[[BeyondCorp]]・[[Binary Authorization]]・[[Certificate Transparency]]・[[Tricorder]]・[[ClusterFuzz]]・[[Project Shield]]・[[DiRT]] など——を具体名で開示しながら論じる点にある。ただし著者陣は、本書が推奨する戦略の一部は組織に存在しないインフラ支援を前提とすること、文化面のベストプラクティスはデータで裏づけられないことを明示的に断っている。(Source: [[Building Secure and Reliable Systems]] Preface) 読み方の指針として、第 1 章と第 2 章を読んだあとは関心のある章から読んでよいとされ、各章は問題設定・ライフサイクル上の適用時期・信頼性とセキュリティの交差点/トレードオフを示す囲みで始まる。本書は料理本(cookbook)ではなく、読者が自組織のリスク環境に合わせて解を作り替えることを前提とする。(Source: [[Building Secure and Reliable Systems]] Preface §How to Read This Book) ## 関連 - 姉妹書: [[SRE Book]] / [[SRE Workbook]] - 編者: [[Heather Adkins]] / [[Betsy Beyer]] / [[Paul Blankinship]] / [[Piotr Lewandowski]] / [[Ana Oprea]] / [[Adam Stubblefield]] - 組織: [[Google]] / [[O'Reilly Media]] - 主要概念: [[信頼性とセキュリティの交差]] / [[最小権限設計]] / [[理解容易性のための設計]] / [[多層防御]] / [[影響範囲の制御]] / [[復旧のための設計]] / [[ソフトウェアサプライチェーンセキュリティ]] / [[セキュリティと信頼性の文化]] ## 出典 - Heather Adkins, Betsy Beyer, Paul Blankinship, Piotr Lewandowski, Ana Oprea, Adam Stubblefield (eds.), *Building Secure and Reliable Systems: Best Practices for Designing, Implementing, and Maintaining Systems*, O'Reilly Media, 2020. https://google.github.io/building-secure-and-reliable-systems/raw/toc.html