# Sun Yat-sen University 広州の研究大学(SYSU、中山大学)。[[Pengfei Chen]] グループが AIOps・マイクロサービス信頼性・根本原因分析の研究を牽引する。(Source: [[@2026__arXiv__Cloud-OpsBench - A Reproducible Benchmark for Agentic Root Cause Analysis in Cloud Systems]]) - [[@2022__ISSRE__Going through the Life Cycle of Faults in Clouds - Guidelines on Fault Handling]] では、[[Xiaoyun Li]]・[[Guangba Yu]]・[[Pengfei Chen]](責任著者)・[[Hongyang Chen]] が SYSU 側として参加し、[[Bizseer]] の [[Zhekang Chen]] と共同で 354 件のポストモーテムを分析したクラウド障害ライフサイクル研究を発表した。 - [[Cloud-OpsBench]] では SYSU 側として Zirui Wang・[[Pengfei Chen]] が参加し、[[The Chinese University of Hong Kong]] と共同で開発した。 - 本論文の参考文献からは、同グループ由来の AIOps 研究が広範に連なる(MicroRank・Nezha・ChangeRCA・SwissLog・MicroSketch・FaaSRCA など、多くで [[Guangba Yu]]・[[Pengfei Chen]] が共著)。 - [[AlertGuardian]](ASE 2025)では [[Guangba Yu]]・Genting Mai・[[Pengfei Chen]](corresponding author)が SYSU 側として参加し、[[Tencent]] と共同でアラートライフサイクル管理フレームワークを開発した。(Source: [[@2025__ASE__AlertGuardian - Intelligent Alert Life-Cycle Management for Large-scale Cloud Systems]]) - [[eACGM]](IWQoS 2025)は [[Ruilin Xu]]・[[Zongxuan Xie]]・[[Pengfei Chen]] による SYSU 単独の研究で、eBPF + libnvml + GMM のフルスタック非侵入 ML システム監視・異常検知フレームワークを開発した。(Source: [[@2025__IWQoS__eACGM - Non-instrumented Performance Tracing and Anomaly Detection towards Machine Learning Systems]]) - [[@2026__arXiv__TELLER - Non-intrusive Cross-Layer Root-Cause Analysis for LLM Inference|TELLER]](ASE '26)は [[Ruilin Xu]]・[[Junyi Li]]・[[Pengfei Chen]]・[[Zongxuan Xie]] による SYSU 単独の研究で、eACGM と同じ著者陣が NVTX/CUPTI ベースの非侵入トレーシングを LLM 推論の異常検知(訓練クラスタ監視)からリクエストレベルの根本原因分析(推論サービング、trace–log マルチモーダル診断)へ発展させた。(Source: [[@2026__arXiv__TELLER - Non-intrusive Cross-Layer Root-Cause Analysis for LLM Inference]]) - [[L4]](FSE 2025)では [[Zhuangbin Chen]]([[Sun Yat-sen University]] Zhuhai)が SYSU 側として参加し、[[The Chinese University of Hong Kong]]・[[Huawei Cloud]] と共同で LLM 訓練障害のログ解析診断フレームワークを開発した。(Source: [[@2025__ESEC-FSE__L4 - Diagnosing Large-scale LLM Training Failures via Automated Log Analysis]]) - [[Tracezip]](ISSTA 2025)は [[Zhuangbin Chen]](筆頭著者)・[[Zibin Zheng]](corresponding author)による SYSU School of Software Engineering(珠海)の研究で、分散トレースのオンライン圧縮システムを [[OpenTelemetry]] Collector 内に実装した。(Source: [[@2025__ISSTA__Tracezip - Efficient Distributed Tracing via Trace Compression]]) - [[Mint]](ASPLOS 2025)では Haiyu Huang・[[Guangba Yu]]・Zilong He・Yilun Wang・[[Pengfei Chen]](corresponding author)が SYSU 側として参加し、[[Alibaba Group]] と共同でコスト効率的な分散トレーシングフレームワークを開発した。(Source: [[@2025__ASPLOS__Mint - Cost-Efficient Tracing with All Requests Collection via Commonality and Variability Analysis]]) - [[LogReducer]](ICSE 2023)では [[Guangba Yu]](筆頭著者)・[[Pengfei Chen]](corresponding author)・Zibin Zheng が SYSU 側として参加し、[[Tencent]] と共同で eBPF ベースのログホットスポット削減フレームワークを開発した。[[WeChat]] 本番に導入しストレージ 39.08% 削減を達成。(Source: [[@2023__ICSE__LogReducer - Identify and Reduce Log Hotspots in Kernel on the Fly]]) - [[@2022__ICWS__TS-InvarNet - Anomaly Detection and Localization based on Tempo-spatial KPI Invariants in Distributed Services|TS-InvarNet]](ICWS 2022)は Zijun Hu・[[Pengfei Chen]](責任著者)・[[Guangba Yu]]・Zilong He・Xiaoyun Li による SYSU 単独の研究で、KPI 不変条件ネットワーク(SARIMAX + HDBSCAN)に基づく解釈可能な異常検知・根本原因箇所特定フレームワークを提案した。モデルサイズ 292 KB、最良 F1 でベースライン比最大 27% 改善。(Source: [[@2022__ICWS__TS-InvarNet - Anomaly Detection and Localization based on Tempo-spatial KPI Invariants in Distributed Services]]) - [[@2023__arXiv__Eadro - An End-to-End Troubleshooting Framework for Microservices on Multi-source Data|Eadro]](arXiv 2302.05092, 2023)では [[Yuxin Su]] が SYSU ソフトウェア工学院として参加し、[[The Chinese University of Hong Kong]] グループの [[Cheryl Lee]]・[[Tianyi Yang]]・[[Zhuangbin Chen]]・[[Michael R. Lyu]] と共同でマルチソースデータ(トレース・ログ・KPI)を用いたマイクロサービスのエンドツーエンドトラブルシューティングフレームワークを構築した。CUHK 単独でなく SYSU が連携する産学共同研究体制の典型例。(Source: [[@2023__arXiv__Eadro - An End-to-End Troubleshooting Framework for Microservices on Multi-source Data]]) - [[@2025__NeurIPS__DynaPipe - Dynamic Layer Redistribution for Efficient Serving of LLMs with Pipeline Parallelism|DynaPipe]](NeurIPS 2025)は [[Hongxin Xu]]・[[Tianyu Guo]](共同筆頭著者)・[[Xianwei Zhang]](責任著者)による CSE 学部単独の研究で、AIOps・信頼性グループ([[Pengfei Chen]] 系列)とは異なる LLM システム研究グループが、パイプライン並列 LLM サービングの動的層再配分手法を提案した。SYSU CSE には AIOps/信頼性と LLM 推論システムという複数の独立した研究系列が併存することを示す。(Source: [[@2025__NeurIPS__DynaPipe - Dynamic Layer Redistribution for Efficient Serving of LLMs with Pipeline Parallelism]]) - [[@2022__ESEC FSE__Constructing Large-Scale Real-World Benchmark Datasets for AIOps]](ESEC/FSE 2022 Industry Track)では [[Pengfei Chen]] が SYSU 側(³)として参加し、[[Tsinghua University]]([[Zeyan Li]]・[[Nengwen Zhao]]・[[Dan Pei]])・[[Nankai University]]([[Shenglin Zhang]]・[[Yongqian Sun]])・[[Microsoft Research]]([[Minghua Ma]])と共同で、KPI 異常検知・多次元根本原因箇所特定・障害発見/診断の 3 公開データセットと年次 AIOps アルゴリズムコンペティションを紹介した。(Source: [[@2022__ESEC FSE__Constructing Large-Scale Real-World Benchmark Datasets for AIOps]]) - [[OpsLLM]]([[@2026__arXiv__OpsLLM - Construction of Large Language Model for Software Operations with Multi-stage Learning]])は [[Jingkai He]](第一著者)・[[Pengfei Chen]](corresponding author)・[[Chenghui Wu]]・[[Shuang Liang]]・[[Gou Tan]]・[[Chuanfu Zhang]] が SYSU 側として参加し、[[Alibaba Cloud]]([[Ye Li]]・[[Xidao Wen]]・[[Fang Situ]]・[[Qi Zhou]])と共同で、QA と RCA を統合的にサポートするソフトウェア運用ドメイン特化 LLM を構築した。同グループの AIOps/RCA 研究群(Nezha・FaaSRCA・TS-InvarNet 等)を、伝統的な決定論的アルゴリズムから LLM 自体の事後訓練(SFT + DPRM ベース強化学習)へ拡張する系譜として位置づけられる。(Source: [[@2026__arXiv__OpsLLM - Construction of Large Language Model for Software Operations with Multi-stage Learning]]) [[@2026__ASE__AgentChaos - Chaos Engineering for Agent Systems via Programmatic Fault Injection|AgentChaos]](ASE 2026)は [[Gou Tan]](第一著者)・[[Zilong He]]・[[Qingfu Wu]]・[[Shuai Liang]]・[[Pengfei Chen]](corresponding author)・[[Chuanfu Zhang]] が SYSU 側として参加し、[[Singapore Management University]]([[Zhensu Sun]]・[[Jieke Shi]]・[[Weifeng Sun]]・[[Junda He]]・[[Lwin Khin Shar]]・[[David Lo]])・Monash University(Ting Zhang)と共同で、LLM API層への非侵入的な障害注入によるエージェントシステム頑健性評価フレームワークを提案した。同グループの AIOps/RCA 研究群(Nezha・FaaSRCA・LLMRCA 等)を、決定論的アルゴリズムやポストホックな根本原因分析から、エージェントシステムそのものの信頼性を制御実験で評価するカオスエンジニアリング領域へ広げた研究として位置づけられる。(Source: [[@2026__ASE__AgentChaos - Chaos Engineering for Agent Systems via Programmatic Fault Injection]]) [[@2026__SIGCOMM__AIDA - Accelerating Root Cause Analysis for Multi-Vendor Device Failures with LLM-Powered Reasoning|AIDA]](SIGCOMM 2026)は [[Haoran Xu]](筆頭著者)・[[Xiaoxi Zhang]](corresponding author)・[[Deke Guo]] が SYSU 側として参加し、[[Alibaba Cloud]]([[Xuan Zeng]]・[[Xumiao Zhang]]・[[Yang Lv]]・[[Peng Zhang]]・[[Ennan Zhai]])と共同で、マルチベンダーのネットワーク機器障害向け自動RCAシステム [[AIDA]] を開発した。Haoran Xu は Alibaba Cloud でのインターンシップ期間中に本研究を実施。同グループの AIOps/RCA 研究群がマイクロサービス・クラウドインフラを主対象としてきたのに対し、AIDA は「ネットワーク機器のマルチベンダー故障」という新しいドメインへ RCA を拡張した点で位置づけられる。(Source: [[@2026__SIGCOMM__AIDA - Accelerating Root Cause Analysis for Multi-Vendor Device Failures with LLM-Powered Reasoning]]) ## 関連 - 本ソース: [[@2026__arXiv__OpsLLM - Construction of Large Language Model for Software Operations with Multi-stage Learning]] / [[@2026__arXiv__Cloud-OpsBench - A Reproducible Benchmark for Agentic Root Cause Analysis in Cloud Systems]] / [[@2025__ASE__AlertGuardian - Intelligent Alert Life-Cycle Management for Large-scale Cloud Systems]] / [[@2025__IWQoS__eACGM - Non-instrumented Performance Tracing and Anomaly Detection towards Machine Learning Systems]] / [[@2025__ESEC-FSE__L4 - Diagnosing Large-scale LLM Training Failures via Automated Log Analysis]] / [[@2025__ASPLOS__Mint - Cost-Efficient Tracing with All Requests Collection via Commonality and Variability Analysis]] / [[@2023__ICSE__LogReducer - Identify and Reduce Log Hotspots in Kernel on the Fly]] / [[@2022__ICWS__TS-InvarNet - Anomaly Detection and Localization based on Tempo-spatial KPI Invariants in Distributed Services]] / [[@2023__arXiv__Eadro - An End-to-End Troubleshooting Framework for Microservices on Multi-source Data]] / [[@2025__NeurIPS__DynaPipe - Dynamic Layer Redistribution for Efficient Serving of LLMs with Pipeline Parallelism]] / [[@2026__ASE__AgentChaos - Chaos Engineering for Agent Systems via Programmatic Fault Injection]] / [[@2026__SIGCOMM__AIDA - Accelerating Root Cause Analysis for Multi-Vendor Device Failures with LLM-Powered Reasoning]] - 所属研究者: [[Pengfei Chen]] / [[Guangba Yu]] / [[Ruilin Xu]] / [[Zongxuan Xie]] / [[Junyi Li]] / [[Zhuangbin Chen]] / [[Yuxin Su]] / [[Hongxin Xu]] / [[Tianyu Guo]] / [[Xianwei Zhang]] / [[Jingkai He]] / [[Gou Tan]] / [[Chenghui Wu]] / [[Shuang Liang]] / [[Chuanfu Zhang]] / [[Zilong He]] / [[Qingfu Wu]] / [[Shuai Liang]] / [[Haoran Xu]] / [[Xiaoxi Zhang]] / [[Deke Guo]] - 共同研究先: [[The Chinese University of Hong Kong]] / [[Tencent]] / [[Huawei Cloud]] / [[Alibaba Cloud]] / [[Singapore Management University]] - 関連プロダクト: [[OpsLLM]] / [[Cloud-OpsBench]] / [[AlertGuardian]] / [[eACGM]] / [[L4]] / [[Tracezip]] / [[Mint]] / [[LogReducer]] / AgentChaos / [[AIDA]] - 関連概念: [[根本原因分析]] / [[AIOps]] / [[異常検知]] / [[ログ解析]] / [[分散トレーシング]] / [[パイプライン並列]] / [[障害注入]] / [[カオスエンジニアリング]] / [[グラフベースRCA]]