# Pengfei Chen
> [!note] 同名異人
> **Xi'an Jiaotong University で 2016 年に博士号を取得した CauseInfer 第一著者**([[Pengfei Chen (Xi'an Jiaotong University)]]、メール:
[email protected])とは別人。本ページは Sun Yat-sen University の研究者(
[email protected])を記録する。
[[Sun Yat-sen University]] の研究者(メール
[email protected])。AIOps・マイクロサービスの根本原因分析・異常検知を長く牽引する。[[Cloud-OpsBench]] の共著者。(Source: [[@2026__arXiv__Cloud-OpsBench - A Reproducible Benchmark for Agentic Root Cause Analysis in Cloud Systems]])
- 本論文の参考文献では、MicroRank(WWW 2021)・Nezha(ESEC/FSE 2023)・ChangeRCA(FSE 2024)・SwissLog(TDSC 2023)・MicroSketch(ICSOC 2022)・FaaSRCA(ISSRE 2024)・TS-InvarNet(ICWS 2022)など、[[Guangba Yu]] との共著を含む AIOps/RCA 研究が広範に並ぶ。
- **[[FaaSRCA]]**(arXiv 2412.02239、2024年12月; ISSRE 2024 掲載とも言及)では責任著者(
[email protected])を務め、Jin Huang・[[Guangba Yu]]・Yilun Wang・Haiyu Huang・Zilong He と共に、サーバーレスアプリケーション向けライフサイクル全体 RCA フレームワークを提案した。Global Call Graph (GCG) でプラットフォーム側(Kubernetes)とアプリ側(関数インスタンス)を統合し、GAT ベースグラフオートエンコーダによる教師なし手法で、Serverless TrainTicket・ML Workflow において平均 HR@k 91.54%・NDCG@k 94.62% を達成し最良ベースライン比 21.25% 改善した。(Source: [[@2024__arXiv__FaaSRCA - Full Lifecycle Root Cause Analysis for Serverless Applications]])
- [[@2022__ICWS__TS-InvarNet - Anomaly Detection and Localization based on Tempo-spatial KPI Invariants in Distributed Services|TS-InvarNet]](ICWS 2022)では責任著者(
[email protected])を務め、Zijun Hu・[[Guangba Yu]]・Zilong He・Xiaoyun Li と共に KPI 不変条件ネットワークに基づく解釈可能な異常検知・箇所特定フレームワークを提案した。SARIMAX モデルと HDBSCAN を組み合わせてシステムトポロジ不要での根本原因 KPI 特定を実現し、最良 F1 スコアでベースラインを最大 27% 上回った。(Source: [[@2022__ICWS__TS-InvarNet - Anomaly Detection and Localization based on Tempo-spatial KPI Invariants in Distributed Services]])
- [[AlertGuardian]](ASE 2025)では corresponding author を務め(
[email protected])、[[Guangba Yu]] らと [[Tencent]] 共同のアラートライフサイクル管理フレームワークを主導した。(Source: [[@2025__ASE__AlertGuardian - Intelligent Alert Life-Cycle Management for Large-scale Cloud Systems]])
- [[eACGM]](IWQoS 2025)でも責任著者(
[email protected])。AIOps/RCA に加え、eBPF + libnvml + GMM による AI/ML システムの非侵入なフルスタック監視・異常検知へ研究を広げる。(Source: [[@2025__IWQoS__eACGM - Non-instrumented Performance Tracing and Anomaly Detection towards Machine Learning Systems]])
- [[Mint]](ASPLOS 2025)では corresponding author を務め、[[Alibaba Group]] と共同でコスト効率的な分散トレーシングフレームワークを開発した。「共通性 + 可変性」パラダイムにより全リクエストを捕捉しつつストレージオーバーヘッドを平均 2.7% に削減する。(Source: [[@2025__ASPLOS__Mint - Cost-Efficient Tracing with All Requests Collection via Commonality and Variability Analysis]])
- [[LogReducer]](ICSE 2023)では corresponding author を務め、[[Guangba Yu]] らと [[Tencent]] 共同で eBPF ベースのログホットスポット削減フレームワークを開発した。[[WeChat]] 本番で 2 か月運用しストレージ 39.08% 削減を達成。(Source: [[@2023__ICSE__LogReducer - Identify and Reduce Log Hotspots in Kernel on the Fly]])
- [[TraStrainer]](ESEC/FSE 2024)では corresponding author を務め、[[Huawei Technologies]] と共同でシステムランタイム状態を考慮した適応的トレースサンプリング手法を開発した。下流 RCA の Top-1 精度を平均 32.63% 向上。(Source: [[@2024__FSE__TraStrainer - Adaptive Sampling for Distributed Traces with System Runtime State]])
- [[@2026__CoNEXT__ChainScope - Balancing Accuracy and Overhead in Non-intrusive Distributed Tracing of Microservices|ChainScope]](ACM CoNEXT 2026)では corresponding author として[[Huawei Technologies]] と共同研究。eBPF によるカーネル内コンテキスト伝搬と IP レベルタギングで非侵襲・高精度・低オーバーヘッドの分散トレーシングを実現し、既存非侵襲手法に対して精度 2.2×・スループット 1.6× 向上を達成した。(Source: [[@2026__CoNEXT__ChainScope - Balancing Accuracy and Overhead in Non-intrusive Distributed Tracing of Microservices]])
- **CloudRanger**(CCGrid 2018)では [[IBM Research China]] 所属として [[Ping Wang]]([[Peking University]])・Jingmin Xu らと共同著者。PC アルゴリズムによるデータ駆動型インパクトグラフ構築と 2 次ランダムウォークを組み合わせたクラウドネイティブ向け根本原因特定フレームワークを提案。IBM Bluemix 本番でトポロジ情報なしに AC@1 59.4%(MonitorRank: 25.4%)を達成。(Source: [[@2018__CCGrid__CloudRanger - Root Cause Identification for Cloud Native Systems]])
- **TraceRank**(JSEP 2021)では責任著者(
[email protected])として [[Guangba Yu]]・[[Zicheng Huang]] と共に、非集計エンドツーエンドトレースを活用したマイクロサービス根本原因分析システムを開発した。スペクトル解析と PageRank ベースのランダムウォークを組み合わせて異常サービスを特定し、Precision 90%・Recall 86% を達成。最良ベースライン比最大 10% の改善を実現した。(Source: [[@2021__JSEP__TraceRank - Abnormal service localization with dis-aggregated end-to-end tracing data in cloud native systems]])
- [[@2018__ICSOC__Microscope - Pinpoint Performance Issues with Causal Graphs in Micro-service Environments|Microscope]](ICSOC 2018)では [[Jinjin Lin]]・[[Zibin Zheng]] と共同で、ソースコード計装なしにネットワークシステムコール傍受と並列化 PC アルゴリズムを組み合わせてマイクロサービス因果グラフを構築し根本原因を特定するシステムを提案した。Sock-shop 評価で PR@1 88% を達成し CauseInfer・CloudRanger・Roots を上回った。(Source: [[@2018__ICSOC__Microscope - Pinpoint Performance Issues with Causal Graphs in Micro-service Environments]])
- **[[@2023__ESEC-FSE__Nezha - Interpretable Fine-Grained Root Causes Analysis for Microservices on Multi-modal Observability Data|Nezha]]**(ESEC/FSE 2023)では責任著者(
[email protected])を務め、[[Guangba Yu]](筆頭著者)・Yufeng Li・Hongyang Chen・Xiaoyun Li・[[Zibin Zheng]] と共に、マルチモーダルオブザーバビリティデータ(メトリクス・トレース・ログ)を均質なイベント表現に統合しイベントグラフを構築することで、コード領域またはリソースタイプレベルの細粒度かつ解釈可能な根本原因分析フレームワークを提案した。構築フェーズの障害フリーイベントパターンと本番フェーズのパターンを比較する完全教師なし手法で、OnlineBoutique と TrainTicket での評価においてサービス内レベル top-1 精度 89.77% を達成し、先行マルチモーダル手法 PDiagnose を TrainTicket で $A^S@5$ 75.56 ポイント上回った。実装: https://github.com/IntelligentDDS/Nezha (Source: [[@2023__ESEC-FSE__Nezha - Interpretable Fine-Grained Root Causes Analysis for Microservices on Multi-modal Observability Data]])
ACM TOSEM Vol. 35 No. 1(2025 年 12 月)に掲載されたサーベイ "A Survey on Failure Analysis and Fault Injection in AI Systems" では corresponding author を務める。[[Guangba Yu]]・[[Gou Tan]]・Haojia Huang・Zhenyu Zhang と共に 142 本の研究を体系的に整理し、AI システムの障害分析(FA)と障害注入(FI)を 6 層タクソノミ(データ収集・前処理・障害検知・根本原因分析・障害注入・ベンチマーク)に体系化した。共著者に [[Roberto Natella]]([[University of Naples Federico II]])・[[Zibin Zheng]]・[[Michael R. Lyu]] が名を連ねる。(Source: [[@2025__TOSEM__A Survey on Failure Analysis and Fault Injection in AI Systems]])
## 関連
- 本ソース: [[@2021__JSEP__TraceRank - Abnormal service localization with dis-aggregated end-to-end tracing data in cloud native systems]] / [[@2018__CCGrid__CloudRanger - Root Cause Identification for Cloud Native Systems]] / [[@2018__ICSOC__Microscope - Pinpoint Performance Issues with Causal Graphs in Micro-service Environments]] / [[@2023__ESEC-FSE__Nezha - Interpretable Fine-Grained Root Causes Analysis for Microservices on Multi-modal Observability Data]] / [[@2024__FSE__TraStrainer - Adaptive Sampling for Distributed Traces with System Runtime State]] / [[@2026__arXiv__Cloud-OpsBench - A Reproducible Benchmark for Agentic Root Cause Analysis in Cloud Systems]] / [[@2025__ASE__AlertGuardian - Intelligent Alert Life-Cycle Management for Large-scale Cloud Systems]] / [[@2025__IWQoS__eACGM - Non-instrumented Performance Tracing and Anomaly Detection towards Machine Learning Systems]] / [[@2025__ASPLOS__Mint - Cost-Efficient Tracing with All Requests Collection via Commonality and Variability Analysis]] / [[@2023__ICSE__LogReducer - Identify and Reduce Log Hotspots in Kernel on the Fly]] / [[@2026__CoNEXT__ChainScope - Balancing Accuracy and Overhead in Non-intrusive Distributed Tracing of Microservices]] / [[@2022__ICWS__TS-InvarNet - Anomaly Detection and Localization based on Tempo-spatial KPI Invariants in Distributed Services]] / [[@2024__arXiv__FaaSRCA - Full Lifecycle Root Cause Analysis for Serverless Applications]] / [[@2025__TOSEM__A Survey on Failure Analysis and Fault Injection in AI Systems]]
- 所属: [[Sun Yat-sen University]]
- 共同研究者: [[Guangba Yu]] / [[Michael R. Lyu]] / [[Roberto Natella]]
- 関連プロダクト: [[Cloud-OpsBench]] / [[AlertGuardian]] / [[Mint]] / [[LogReducer]] / [[TraStrainer]] / [[@2023__ESEC-FSE__Nezha - Interpretable Fine-Grained Root Causes Analysis for Microservices on Multi-modal Observability Data|Nezha]] / [[FaaSRCA]]
- 関連概念: [[根本原因分析]] / [[マルチモーダル障害診断]] / [[AIOps]] / [[異常検知]] / [[分散トレーシング]] / [[ログ解析]] / [[eBPF]] / [[サーバーレスRCA]] / [[障害注入]] / [[障害分析]]