# Dongmei Zhang
[[@2013__ASE__Software Analytics for Incident Management of Online Services - An Experience Report]](ASE 2013)の共著者として、SAS(Service Analysis Studio)の開発を [[Jian-Guang Lou]] らと共同実施した。
Microsoft(中国)所属のシニア研究者。StepFlyの共著者として、Microsoft ResearchにおけるAIOps・インシデント管理研究のリーダーシップを担う。TaskWeaver、AllHands等の複数の研究を共同主導している。[[DeepIP]] 論文([[@2020__ASE__How Incidental are the Incidents - Characterizing and Prioritizing Incidents for Large-Scale Online Service Systems]])にも共著者として参加し、Microsoft の 18 オンラインサービス横断のインシデント優先順位付け研究を支援した。
[[@2019__WWW__Outage Prediction and Diagnosis for Cloud Service Systems|AirAlert]](WWW2019)にも Microsoft Research Beijing 所属のシニア共著者として参加し、Microsoft クラウドのアウテージ予測・診断研究を支援した。Microsoft Research の AIOps 研究系譜(iDice ICSE2016・Log Clustering ICSE2016・AirAlert WWW2019・Gandalf NSDI2020・DeepIP ASE2020 等)の継続的な共著者。
COMET 論文([[@2024__ISSRE__Large Language Models Can Provide Accurate and Interpretable Incident Triage]]、ISSRE 2024)の共著者として、LLM キーワード抽出によるインシデントトリアージの Microsoft 本番展開研究に参加。(Source: [[@2024__ISSRE__Large Language Models Can Provide Accurate and Interpretable Incident Triage]])
Aloha 論文([[@2026__FSE Companion__Aloha - Localizing Batch Failures in Large-scale Cloud Systems via Contrast Analysis and Human-in-the-Loop Agent]]、FSE Companion '26)の共著者([[Microsoft]])。[[Shenglin Zhang]](筆頭)・[[Yujia Wu]]・[[Jinghuan Ren]]・[[Wenwei Gu]]([[Nankai University]])、責任著者 [[Yongqian Sun]]([[Nankai University]])、[[Chaoyun Zhang]]・[[Liqun Li]]・[[Qingwei Lin]]・[[Saravan Rajmohan]]・[[Chetan Bansal]]・[[Minghua Ma]]([[Microsoft]])との共同で、対照分析ベースの異常箇所特定を human-in-the-loop エージェントでオペレーショナル化するフレームワーク Aloha を提案した。(Source: [[@2026__FSE Companion__Aloha - Localizing Batch Failures in Large-scale Cloud Systems via Contrast Analysis and Human-in-the-Loop Agent]])
## 関連
- ソース: [[@2013__ASE__Software Analytics for Incident Management of Online Services - An Experience Report]] / [[@2019__WWW__Outage Prediction and Diagnosis for Cloud Service Systems]] / [[@2025__arXiv__StepFly - Agentic Troubleshooting Guide Automation for Incident Diagnosis]] / [[@2021__ISSRE__How Long Will it Take to Mitigate this Incident for Online Service Systems]] / [[@2020__ASE__How Incidental are the Incidents - Characterizing and Prioritizing Incidents for Large-Scale Online Service Systems]] / [[@2025__KDD__Large Language Models can Deliver Accurate and Interpretable Time Series Anomaly Detection]] / [[@2024__ISSRE__Large Language Models Can Provide Accurate and Interpretable Incident Triage]]
- 所属: [[Microsoft]]