# Peng Huang [[University of Michigan]] [[OrderLab]] の PI(Principal Investigator)。DL 訓練システムの信頼性、分散システムのサイレント障害検知、プログラム不変条件推論を研究する。 [[Microsoft Research]] 在籍時(2017年頃)に [[Lidong Zhou]] らと共著で "Gray Failure: The Achilles' Heel of Cloud-Scale Systems"(HotOS 2017)を発表し、**グレイ障害(differential observability)**という概念を最初に定式化した。[[Johns Hopkins University]] との兼任期間を経て[[University of Michigan]]に移籍。 Gandalf 論文([[@2020__NSDI__Gandalf - An Intelligent, End-To-End Analytics Service for Safe Deployment in Large-Scale Cloud Infrastructure]]、NSDI 2020)の共著者。所属は [[Johns Hopkins University]] と記載されており、Microsoft Research(2017年の Gray Failure 論文時点)から Johns Hopkins University(2020年、本論文)を経て University of Michigan に移籍したという既知のキャリアパスと時系列的に整合する。[[Ze Li]] ら Microsoft Azure チームとの共同研究として、クラウドロールアウトの安全性判定サービス [[Gandalf]] の研究に参加した。(Source: [[@2020__NSDI__Gandalf - An Intelligent, End-To-End Analytics Service for Safe Deployment in Large-Scale Cloud Infrastructure]]) ## 主な貢献 - [[@2017__HotOS__Gray Failure - The Achilles' Heel of Cloud-Scale Systems]]: 第一著者。Azure 本番インシデント経験から差分可観測性(differential observability)によるグレイ障害の定義を提唱。HotOS 2017。 - [[@2020__NSDI__Gandalf - An Intelligent, End-To-End Analytics Service for Safe Deployment in Large-Scale Cloud Infrastructure]]: 共著者(Johns Hopkins University 所属時)。NSDI 2020。 - [[@2025__OSDI__Training with Confidence - Catching Silent Errors in Deep Learning Training with Automated Proactive Checks]](TrainCheck): シニア著者。OSDI 2025。 ## 関連 - ソース: [[@2017__HotOS__Gray Failure - The Achilles' Heel of Cloud-Scale Systems]] / [[@2020__NSDI__Gandalf - An Intelligent, End-To-End Analytics Service for Safe Deployment in Large-Scale Cloud Infrastructure]] / [[@2025__OSDI__Training with Confidence - Catching Silent Errors in Deep Learning Training with Automated Proactive Checks]] - エンティティ: [[Yuxuan Jiang]] / [[University of Michigan]] / [[OrderLab]] / [[TrainCheck]] / [[Lidong Zhou]] / [[Jacob R. Lorch]] / [[Microsoft Research]] / [[Johns Hopkins University]] / [[Ze Li]] / [[Gandalf]]