# Chetan Bansal [[Microsoft]](Redmond, USA)所属([email protected])。[[FLASH]]([[@2024__MSR__FLASH - A Workflow Automation Agent for Diagnosing Recurring Incidents]])の共著者であり、Microsoft のクラウドインシデントを対象とした LLM 活用 RCA 研究(ICSE 2023)にも参加している(Source: [[@2024__MSR__FLASH - A Workflow Automation Agent for Diagnosing Recurring Incidents]])。AIOpsLab など Microsoft AIOps 研究群における主要な研究者の一人。 ESEC/FSE 2023 の論文 "Detection Is Better Than Cure: A Cloud Incidents Perspective"([[@2023__ESEC-FSE__Detection Is Better Than Cure - A Cloud Incidents Perspective]])の共著者として、Microsoft の 300 超サービスのモニタリングギャップ実証分析にも参加(Source: [[@2023__ESEC-FSE__Detection Is Better Than Cure - A Cloud Incidents Perspective]])。 [[LLMAD]] 論文([[@2025__KDD__Large Language Models can Deliver Accurate and Interpretable Time Series Anomaly Detection]], KDD 2025)の共著者。LLM を直接 TSAD に使う初のフレームワーク。本論文の方向性は本人が共著者として参加した「Detection Is Better Than Cure(ESEC/FSE 2023)」が示した「ミス検知の原因はモニタの欠如」という問題意識の延長として、LLM による低コスト・解釈可能な検知器を提案するものに位置づく。(Source: [[@2025__KDD__Large Language Models can Deliver Accurate and Interpretable Time Series Anomaly Detection]]) さらに ICSE-SEIP 2024 の "Intelligent Monitoring Framework for Cloud Services: A Data-Driven Approach"([[@2024__ICSE-SEIP__Intelligent Monitoring Framework for Cloud Services - A Data-Driven Approach]])の共著者として、791 本番サービス・30,920 モニタのオントロジー構築とプロトタイプ学習ベースのモニタ推奨フレームワークの研究に参加(Source: [[@2024__ICSE-SEIP__Intelligent Monitoring Framework for Cloud Services - A Data-Driven Approach]])。 その続編にあたる FSE 2026 industry 論文 "Attention Enhanced Entity Recommendation for Intelligent Monitoring in Cloud Systems"([[@2026__FSE__Attention Enhanced Entity Recommendation for Intelligent Monitoring in Cloud Systems]])の共著者として、ディメンション部分集合推薦のためのヘテロジニアスグラフニューラルネットワーク DiRecGNN(本番モニタ 18,291・メトリクス 4,623・ディメンション 8,356 で HR@1 +55.8%・MRR +43.1% 改善)の研究に参加。 FSE 2024 "X-lifecycle Learning for Cloud Incident Management using LLMs"([[@2024__FSE__X-lifecycle Learning for Cloud Incident Management using LLMs]])の共著者として、X-lifecycle データを LLM プロンプトに補完した RCA 精度改善と監視オントロジー識別の研究に参加(Source: [[@2024__FSE__X-lifecycle Learning for Cloud Incident Management using LLMs]])。 AgentTether 論文([[@2026__arXiv__AgentTether - Graph-Guided Diagnosis and Runtime Intervention for Reliable LLM Agent Operations]], arXiv 2026-07)の共著者。[[Chenyu Zhao]]・[[Shenglin Zhang]]([[Nankai University]])らとの共同で、LLM エージェントの失敗トレースをグラフ表現で局所診断し実行時に修復する AgentTether を提案。これまでのクラウドインシデント RCA・モニタリング研究(FLASH・eARCO 等)の対象がインフラ/サービスの障害だったのに対し、本論文は**エージェント自身の意思決定軌跡**を診断対象とする点で新しい。(Source: [[@2026__arXiv__AgentTether - Graph-Guided Diagnosis and Runtime Intervention for Reliable LLM Agent Operations]]) PROBE 論文([[@2026__arXiv__Debugging the Debuggers - Failure-Anchored Structured Recovery for Software Engineering Agents]], arXiv 2605.08717)の共著者。[[Chenyu Zhao]]・[[Shenglin Zhang]]([[Nankai University]])・[[Dan Pei]]([[Tsinghua University]])・[[Saravan Rajmohan]]・[[Minghua Ma]] との共同で、AgentTether に先行する失敗起点構造化回復フレームワークを提案。ソフトウェアエンジニアリングエージェントの失敗した run のテレメトリを構造化証拠・構造化診断・範囲限定ガイダンスへ変換し、Microsoft IcM のサービス診断エージェントワークフローへ非侵襲的な side channel として統合するプロトタイプ検証も行った。(Source: [[@2026__arXiv__Debugging the Debuggers - Failure-Anchored Structured Recovery for Software Engineering Agents]]) COMET 論文([[@2024__ISSRE__Large Language Models Can Provide Accurate and Interpretable Incident Triage]]、ISSRE 2024)の共著者。LLM キーワード抽出によるインシデントトリアージシステムの Microsoft 本番展開研究に参加。(Source: [[@2024__ISSRE__Large Language Models Can Provide Accurate and Interpretable Incident Triage]]) Build-bench 論文([[@2026__nkcs.iops.ai__Can Language Models Go Beyond Coding - Assessing the Capability of Language Models to Build Real-World Systems]], nkcs.iops.ai 2026-05)の共著者([[Microsoft]] Redmond, USA)。[[Chenyu Zhao]]・[[Shenglin Zhang]]・[[Yongqian Sun]]([[Nankai University]])、[[Weilin Jin]]([[Peking University]])、[[Dan Pei]]([[Tsinghua University]])、[[Chaoyun Zhang]]・[[Qingwei Lin]]・[[Saravan Rajmohan]]・[[Minghua Ma]]との共同で、クロス ISA(x86_64/aarch64)ビルド失敗の LLM 修復能力を評価する初のベンチマーク Build-bench を提案した。AgentTether・PROBE(いずれも同じ Nankai チームとの共著)がエージェント自身の実行トレース診断を扱うのに対し、本論文はソフトウェアビルド・パッケージング領域での LLM 自律修復能力を扱う。(Source: [[@2026__nkcs.iops.ai__Can Language Models Go Beyond Coding - Assessing the Capability of Language Models to Build Real-World Systems]]) Aloha 論文([[@2026__FSE Companion__Aloha - Localizing Batch Failures in Large-scale Cloud Systems via Contrast Analysis and Human-in-the-Loop Agent]]、FSE Companion '26)の共著者([[Microsoft]])。[[Shenglin Zhang]](筆頭)・[[Yujia Wu]]・[[Jinghuan Ren]]・[[Wenwei Gu]]([[Nankai University]])、責任著者 [[Yongqian Sun]]([[Nankai University]])、[[Chaoyun Zhang]]・[[Liqun Li]]・[[Qingwei Lin]]・[[Dongmei Zhang]]・[[Saravan Rajmohan]]・[[Minghua Ma]]([[Microsoft]])との共同で、対照分析ベースの異常箇所特定を human-in-the-loop エージェントでオペレーショナル化するフレームワーク Aloha を提案した。(Source: [[@2026__FSE Companion__Aloha - Localizing Batch Failures in Large-scale Cloud Systems via Contrast Analysis and Human-in-the-Loop Agent]]) TSGen 論文([[@2026__FSE Companion__TSGen - Automated Troubleshooting Guide Generation]]、FSE Companion '26)の共著者。[[Yi Xiao]]・[[Hongyu Zhang]]([[Chongqing University]])、[[Daniel Genkin]]・[[Chaoyun Zhang]]・[[Rujia Wang]]・[[Bhala Ranganathan]]・[[Saravan Rajmohan]]・[[Minghua Ma]](corresponding author)との共同で、過去インシデントレポートから LLM で構造化 TSG をゼロから自動生成するパイプライン TSGen を提案。[[FLASH]] が「既存 TSG をどう実行するか」を扱うのに対し、TSGen は「TSG が存在しない/古い状態からどう生成するか」という一段上流の課題を扱う。(Source: [[@2026__FSE Companion__TSGen - Automated Troubleshooting Guide Generation]]) Comfey 論文([[@2026__FSE Companion__An Agentic Framework for Triaging Incidents in Production Cloud Infrastructure]]、FSE Companion '26)の共著者。[[Minghua Ma]]・[[Ze Li]]・[[Murali Chintalapati]] らとの共同で、Microsoft Azure 本番環境向けの分散型エージェント型インシデントトリアージフレームワーク [[Comfey]] の研究に参加。本人が共著者である COMET 論文を先行の Azure 本番システムとして直接比較しており、Comfey は COMET 比でトリアージ精度 +7.55%・トリアージ時間 4.38 倍高速化・緩和時間 2.91 倍高速化を達成した。COMET(中央集権的 LLM キーワード抽出)→Comfey(分散型・チームスコープド)というインシデントトリアージ研究の系譜に一貫して関与している。(Source: [[@2026__FSE Companion__An Agentic Framework for Triaging Incidents in Production Cloud Infrastructure]] §4.8) ## 関連 - ソース: [[@2026__FSE Companion__An Agentic Framework for Triaging Incidents in Production Cloud Infrastructure]] / [[@2026__FSE Companion__TSGen - Automated Troubleshooting Guide Generation]] / [[@2022__SoCC__How to Fight Production Incidents]] / [[@2023__ESEC-FSE__Detection Is Better Than Cure - A Cloud Incidents Perspective]] / [[@2024__MSR__FLASH - A Workflow Automation Agent for Diagnosing Recurring Incidents]] / [[@2024__ICSE-SEIP__Intelligent Monitoring Framework for Cloud Services - A Data-Driven Approach]] / [[@2024__FSE__X-lifecycle Learning for Cloud Incident Management using LLMs]] / [[@2026__FSE__Attention Enhanced Entity Recommendation for Intelligent Monitoring in Cloud Systems]] / [[@2026__TSC__LLM-Enhanced Failure Localization in Microservices - Integrating Multi-Modal Data and Expert Interpretation]] / [[@2026__arXiv__AgentTether - Graph-Guided Diagnosis and Runtime Intervention for Reliable LLM Agent Operations]] / [[@2026__arXiv__Debugging the Debuggers - Failure-Anchored Structured Recovery for Software Engineering Agents]] / [[@2026__nkcs.iops.ai__Can Language Models Go Beyond Coding - Assessing the Capability of Language Models to Build Real-World Systems]] - 所属: [[Microsoft]] - 共著者: [[Supriyo Ghosh]] / [[Suman Nath]] / [[Xuchao Zhang]] / [[Minghua Ma]] / [[Saravan Rajmohan]] / [[Vaibhav Ganatra]] / [[Yu Kang]] / [[Jonathan Mace]] / [[Pooja Srinivas]] / [[Fiza Husain]] / [[Anjaly Parayil]] / [[Drishti Goel]] / [[Aditya Singh]] / [[Ayush Choure]] / [[Anson Bastos]] / [[Rujia Wang]] / [[Chenyu Zhao]] / [[Shenglin Zhang]] / [[Weilin Jin]]