# Yongqian Sun [[Nankai University]] Department of Software Engineering の助教。AIOps(異常検知・障害予測・変更管理)を中心に [[Shenglin Zhang]] と継続的に共同研究する。[[OptProphet]] 論文([[@2025__APNET__Forewarned is Forearmed - Joint Prediction and Classification of Optical Transceiver Failures in Large-Scale LLM Training Clusters]], APNet 2025)の共著者。[[SCELM]] 論文([[@2025__FSE Companion__A Multimodal Intelligent Change Assessment Framework for Microservice Systems Based on Large Language Models]], FSE Companion '25)の第一著者。また [[SuperAgg]] フレームワーク([[@2024__ISSRE__Exploring Hierarchical Patterns for Alert Aggregation in Supercomputers]], ISSRE 2024)の共著者として [[National University of Defense Technology]](NUDT)との共同研究にも参加している。(Source: [[@2025__FSE Companion__A Multimodal Intelligent Change Assessment Framework for Microservice Systems Based on Large Language Models]]、[[@2024__ISSRE__Exploring Hierarchical Patterns for Alert Aggregation in Supercomputers]]) [[@2023__TSC__LogKG - Log Failure Diagnosis through Knowledge Graph|LogKG]] 論文([[@2023__TSC__LogKG - Log Failure Diagnosis through Knowledge Graph]], IEEE Transactions on Services Computing 2023)の共著者。[[Yicheng Sui]]・[[Shenglin Zhang]]・[[Dan Pei]] らとともに、ログの複数フィールドを知識グラフで統合する障害診断フレームワーク LogKG の開発に参加した。(Source: [[@2023__TSC__LogKG - Log Failure Diagnosis through Knowledge Graph]]) [[UniDiag]] 論文([[@2024__TSC__No More Data Silos - Unified Microservice Failure Diagnosis With Temporal Knowledge Graph]], IEEE TSC 2024)の corresponding author([email protected])。[[Shenglin Zhang]]・[[Yongxin Zhao]]・[[Dan Pei]] らとの共同で、時系列知識グラフ(TKG)によるマルチモーダル障害診断フレームワーク [[UniDiag]] を提案した。(Source: [[@2024__TSC__No More Data Silos - Unified Microservice Failure Diagnosis With Temporal Knowledge Graph]]) [[LogInsight]] 論文([[@2025__nkcs.iops.ai__Accurate and Interpretable Log-Based Fault Diagnosis using Large Language Models]], 2025)の共同第一著者([[Shiyu Ma]] と並列の \* 印著者)。[[Tong Xiao]]・[[Yongxin Zhao]]・[[Shenglin Zhang]]・[[Dan Pei]] らとともに、LLM を用いた解釈可能なログベース障害診断フレームワーク [[LogInsight]] を開発。FOLS モジュール + LoRA ファインチューニングで 3 データセットすべてで最高性能を達成した。(Source: [[@2025__nkcs.iops.ai__Accurate and Interpretable Log-Based Fault Diagnosis using Large Language Models]]) [[LagRCA]] 論文([[@2026__FSE Companion__Bridging the Delay - Lag-Aware Spatio-Temporal Causal Inference for Microservice Root Cause Analysis]], FSE Companion '26)の共著者。[[Shenglin Zhang]](筆頭)・[[Junhua Kuang]]・[[Yimeng Zhang]]・[[Sibo Xia]]・[[Jintao Feng]]・[[Jingyu Wang]]・[[Wenwei Gu]]([[Nankai University]])、[[Wei Li]]・[[Liping Zhang]]([[Alibaba Group]])、[[Dan Pei]]([[Tsinghua University]])との共同で、マイクロサービス障害の可変時間ラグを明示的にモデル化する遅延認識時空間因果推論フレームワーク LagRCA を提案した。(Source: [[@2026__FSE Companion__Bridging the Delay - Lag-Aware Spatio-Temporal Causal Inference for Microservice Root Cause Analysis]]) AgentTether 論文([[@2026__arXiv__AgentTether - Graph-Guided Diagnosis and Runtime Intervention for Reliable LLM Agent Operations]], arXiv 2026-07)の共著者。[[Chenyu Zhao]]・[[Shenglin Zhang]]・[[Wenwei Gu]]・[[Dan Pei]]・[[Chetan Bansal]]・[[Saravan Rajmohan]]・[[Minghua Ma]]らとの共同で、LLM エージェントの失敗軌跡をグラフ表現で診断し実行時に修復するフレームワークを提案した。(Source: [[@2026__arXiv__AgentTether - Graph-Guided Diagnosis and Runtime Intervention for Reliable LLM Agent Operations]]) [[Build-bench]] 論文([[@2026__nkcs.iops.ai__Can Language Models Go Beyond Coding - Assessing the Capability of Language Models to Build Real-World Systems]], nkcs.iops.ai 2026-05)の共著者。[[Chenyu Zhao]]・[[Shenglin Zhang]](いずれも [[Nankai University]])、[[Weilin Jin]]([[Peking University]])、[[Dan Pei]]([[Tsinghua University]])、[[Chaoyun Zhang]]・[[Qingwei Lin]]・[[Chetan Bansal]]・[[Saravan Rajmohan]]・[[Minghua Ma]]([[Microsoft]])との共同で、クロス ISA(x86_64/aarch64)ビルド失敗の LLM 修復能力を評価する初のベンチマーク Build-bench を提案した。(Source: [[@2026__nkcs.iops.ai__Can Language Models Go Beyond Coding - Assessing the Capability of Language Models to Build Real-World Systems]]) PROBE 論文([[@2026__arXiv__Debugging the Debuggers - Failure-Anchored Structured Recovery for Software Engineering Agents]], arXiv 2605.08717)の共著者。[[Chenyu Zhao]](筆頭)・[[Shenglin Zhang]]・Yihang Lin・[[Wenwei Gu]]・Zhimin Chen・[[Dan Pei]]・[[Chetan Bansal]]・[[Saravan Rajmohan]]・[[Minghua Ma]] との共同で、AgentTether に先行する失敗起点構造化回復フレームワークを提案。SWE-bench・EnterpriseOps-Gym・AIOpsLab 3 設定・257 件の初回未解決ケースで Top-1 診断精度 65.37%・recovery rate 21.79% を達成した。(Source: [[@2026__arXiv__Debugging the Debuggers - Failure-Anchored Structured Recovery for Software Engineering Agents]]) CoTriage 論文([[@2026__nkcs.iops.ai__Collaborative Knowledge Distillation and Reinforcement Learning for Automated Ticket Triage in Large-Scale Production Systems]])の共著者。[[Ruowei Fu]](筆頭)・[[Shenglin Zhang]](責任著者)・[[ByteDance]] の STE チームと共同で、知識蒸留 + 自己強化 + DPO によるチケットトリアージフレームワークを提案した。先行研究 [[OncallX]] 論文([[@2025__ASE__LLM-Powered Multi-Agent Collaboration for Intelligent Industrial On-Call Automation]])にも共著者として参加しており、同一ドメインで fine-tuning-free 路線(OncallX)と蒸留+RL 路線(CoTriage)の双方に関与している。(Source: [[@2026__nkcs.iops.ai__Collaborative Knowledge Distillation and Reinforcement Learning for Automated Ticket Triage in Large-Scale Production Systems]]) InsightTriage 論文([[@2026__ASE__LLM-Assisted Joint Ticket and Log Analysis for Incident Triage in Intelligent and Connected Vehicles]], ASE '26 投稿版)の共著者。[[Ruowei Fu]](筆頭)・[[Shenglin Zhang]](責任著者)・[[Wenwei Gu]](いずれも [[Nankai University]])、Weiguo Li([[Huawei Technologies|Huawei Inc.]])、[[Dan Pei]]([[Tsinghua University]])との共同で、ICV 向けの LLM 支援ジョイントチケット・ログ分析インシデントトリアージシステムを提案した。(Source: [[@2026__ASE__LLM-Assisted Joint Ticket and Log Analysis for Incident Triage in Intelligent and Connected Vehicles]]) OScope 論文([[@2026__ICSE-SEIP__When LLMs Listen to Experts - Accurate Failure Diagnosis in Operating Systems]], ICSE-SEIP '26)の共著者。[[Yongxin Zhao]](筆頭)・[[Shenglin Zhang]](責任著者)・[[Yuxin Sun]]・[[Wenwei Gu]]([[Nankai University]])、[[Alibaba Group]] の [[Luping Wang]]・[[Li Shi]]・[[Cheng Huang]]・[[Guodong Yang]]・[[Liping Zhang]]、[[Dan Pei]]([[Tsinghua University]])との共同で、OS 障害診断向け LLM フレームワーク [[OScope]] を提案した。マイクロサービス([[UniDiag]] 等)から OS レベルへ研究対象を広げた一本。(Source: [[@2026__ICSE-SEIP__When LLMs Listen to Experts - Accurate Failure Diagnosis in Operating Systems]]) [[PerfScout]] 論文([[@2026__ICSE-SEIP__PerfScout - An Adaptive Workload Generator in Software Performance Testing]], ICSE-SEIP '26)の共著者([[Nankai University]])。[[Shenglin Zhang]](責任著者)を筆頭に、Qingliang Zhang・Xiao Xiong・Mengyao Li・Yimin Zuo(いずれも Nankai University)、[[Xidao Wen]]([[BizSeer]])、[[Wenwei Gu]]、Huandong Zhuang・Bowen Deng・Ruiyuan Wan([[Huawei Cloud]])、[[Dan Pei]]([[Tsinghua University]])との共同で、SPOT(極値理論ベースの動的閾値)・ADF/KPSS(局所定常性判定)・PPO(強化学習)を統合した性能テスト向け適応的ワークロード生成フレームワーク [[PerfScout]] を提案した。Huawei Cloud の CodeArts PerfTest に 9 か月間デプロイされ、ブレークポイント特定精度 82% 超・テスト効率改善 90% 近くを実証。同著者グループの先行研究 Auto-PIP(ISSREW 2024)の後継にあたる。(Source: [[@2026__ICSE-SEIP__PerfScout - An Adaptive Workload Generator in Software Performance Testing]]) TADBench 論文([[@2025__TSC__A Comprehensive Benchmark and Empirical Study of Trace Anomaly Detection]], IEEE TSC 2025)の第一著者。[[Minyi Shao]]・[[Xiaohui Nie]]・[[Kaiwen Yang]]・[[Xingda Li]]・[[Bowen Hao]]・[[Shenglin Zhang]]・[[Changhua Pei]]・[[Dongbiao He]]・[[Yanbiao Li]]・[[Dan Pei]] との共同で、5 公開トレースデータセット(TrainTicket・GAIA・AIOps2020/2022/2023)を統一フォーマットに標準化し約 21 万トレースへ人手ラベルを付与、7 種のトレース異常検知アルゴリズムを横断比較する初の包括的ベンチマーク TADBench を提案した。従来の障害診断・根本原因分析(LogInsight・UniDiag 等)がログ・メトリクス中心だったのに対し、本論文はトレースデータそのものの異常検知ベンチマークという新しい研究軸を加える。(Source: [[@2025__TSC__A Comprehensive Benchmark and Empirical Study of Trace Anomaly Detection]]) RefinedEdge 論文([[@2025__TSC__Bridging Edge and Cloud - A Knowledge-Enhanced Framework for Efficient Time Series Anomaly Detection]], IEEE TSC 2025)の corresponding author。[[Shenglin Zhang]](筆頭)・[[Jiacheng Zhang]]・[[Guohua Liu]]([[Alibaba Cloud]])・[[Shiqi Chen]]・[[Chenyu Zhao]]・[[Minghua Ma]]([[Microsoft]])・[[Yutong Chen]]・[[Dan Pei]]([[Tsinghua University]])との共同で、多変量時系列異常検知モデルをエッジ配備可能な水準まで圧縮しつつクラウド訓練モデルに匹敵する精度を達成する知識強化フレームワーク [[RefinedEdge]] を提案した。EdgeNode/SMD/MSL/SMAP の 4 データセットで 0.12M パラメータの個人化学生モデルが F1=0.9588/0.9274/0.8827/0.8580 を達成した。TADBench(トレース異常検知ベンチマーク)に続き、Sun の関与研究にモデル圧縮・知識蒸留・エッジクラウド協調という新しい技術軸が加わった。(Source: [[@2025__TSC__Bridging Edge and Cloud - A Knowledge-Enhanced Framework for Efficient Time Series Anomaly Detection]]) LogSage 論文([[@2025__FCS__From Chaos to Clarity - Log-based Kernel Panic Root Cause Analysis for Large-Scale Cloud Services]], Frontiers of Computer Science 2025)の共著者。[[Tianyu Cui]](筆頭)・[[Shenglin Zhang]](責任著者)・[[Yicheng Sui]]・[[Zeyu Che]](いずれも [[Nankai University]])と [[ByteDance]] との共同で、OS カーネルパニックのログベース RCA フレームワーク LogSage を提案した。GraphSAGE によるグラフ表現学習と能動学習を組み合わせた GARCA モジュールで、ByteDance 本番データを含む 3 データセットで最強ベースライン LogKG を 15.5〜20.3 ポイント上回る F1 を達成した。(Source: [[@2025__FCS__From Chaos to Clarity - Log-based Kernel Panic Root Cause Analysis for Large-Scale Cloud Services]]) Aloha 論文([[@2026__FSE Companion__Aloha - Localizing Batch Failures in Large-scale Cloud Systems via Contrast Analysis and Human-in-the-Loop Agent]]、FSE Companion '26)の責任著者(corresponding author)。[[Shenglin Zhang]](筆頭)・[[Yujia Wu]]・[[Jinghuan Ren]]・[[Wenwei Gu]](いずれも [[Nankai University]])、[[Chaoyun Zhang]]・[[Liqun Li]]・[[Qingwei Lin]]・[[Dongmei Zhang]]・[[Saravan Rajmohan]]・[[Chetan Bansal]]・[[Minghua Ma]]([[Microsoft]])との共同で、対照分析ベースの異常箇所特定を human-in-the-loop エージェントでオペレーショナル化するフレームワーク Aloha を提案した。既存手法 [[CONAN]] に対し ACC@5 で 0.9370 vs 0.6963 と全指標で上回り、Microsoft クラウドの 127 件の実バッチ障害ケースで有効性を実証した。(Source: [[@2026__FSE Companion__Aloha - Localizing Batch Failures in Large-scale Cloud Systems via Contrast Analysis and Human-in-the-Loop Agent]]) OpsMem 論文([[@2026__arXiv__OpsMem - Dual-Memory Reasoning with Cross-Memory Resonance for Failure Diagnosis]], arXiv 2026-07)の第一著者。[[Rongchen Gao]]・[[Yu Luo]]・[[Wenwei Gu]]・[[Shenglin Zhang]](いずれも [[Nankai University]])、[[Qingyi Guo]]・[[Dan Pei]]([[Tsinghua University]])、[[Qiuai Fu]]・[[Yaoliang Wu]]([[Huawei Technologies]])との共同で、短期記憶(STM)と長期記憶(LTM)を cross-memory resonance で結合する失敗診断向けデュアルメモリフレームワーク [[OpsMem]] を提案した。GoS の belief-state 抽象化を STM に踏襲しつつ、[[OpsAgent]] に続く同研究グループの障害診断エージェント系譜に連なる。(Source: [[@2026__arXiv__OpsMem - Dual-Memory Reasoning with Cross-Memory Resonance for Failure Diagnosis]]) ## 関連 - ソース: [[@2025__FCS__From Chaos to Clarity - Log-based Kernel Panic Root Cause Analysis for Large-Scale Cloud Services]] / [[@2026__arXiv__OpsMem - Dual-Memory Reasoning with Cross-Memory Resonance for Failure Diagnosis]] / [[@2025__TSC__Bridging Edge and Cloud - A Knowledge-Enhanced Framework for Efficient Time Series Anomaly Detection]] / [[@2025__TSC__A Comprehensive Benchmark and Empirical Study of Trace Anomaly Detection]] / [[@2026__ASE__LLM-Assisted Joint Ticket and Log Analysis for Incident Triage in Intelligent and Connected Vehicles]] / [[@2024__ASE__ART - A Unified Unsupervised Framework for Incident Management in Microservice Systems]] / [[@2025__APNET__Forewarned is Forearmed - Joint Prediction and Classification of Optical Transceiver Failures in Large-Scale LLM Training Clusters]] / [[@2025__FSE Companion__A Multimodal Intelligent Change Assessment Framework for Microservice Systems Based on Large Knowledge Models]] / [[@2026__ASE__OpsAgent - An Evolving Multi-agent System for Incident Management in Microservices]] / [[@2024__ISSRE__Exploring Hierarchical Patterns for Alert Aggregation in Supercomputers]] / [[@2023__TSC__LogKG - Log Failure Diagnosis through Knowledge Graph]] / [[@2024__TSC__No More Data Silos - Unified Microservice Failure Diagnosis With Temporal Knowledge Graph]] / [[@2025__nkcs.iops.ai__Accurate and Interpretable Log-Based Fault Diagnosis using Large Language Models]] / [[@2026__arXiv__AgentTether - Graph-Guided Diagnosis and Runtime Intervention for Reliable LLM Agent Operations]] / [[@2026__nkcs.iops.ai__Can Language Models Go Beyond Coding - Assessing the Capability of Language Models to Build Real-World Systems]] / [[@2026__arXiv__Debugging the Debuggers - Failure-Anchored Structured Recovery for Software Engineering Agents]] / [[@2026__nkcs.iops.ai__Collaborative Knowledge Distillation and Reinforcement Learning for Automated Ticket Triage in Large-Scale Production Systems]] / [[@2025__ASE__LLM-Powered Multi-Agent Collaboration for Intelligent Industrial On-Call Automation]] / [[@2026__FSE Companion__Bridging the Delay - Lag-Aware Spatio-Temporal Causal Inference for Microservice Root Cause Analysis]] / [[@2026__ICSE-SEIP__When LLMs Listen to Experts - Accurate Failure Diagnosis in Operating Systems]] / [[@2026__ICSE-SEIP__PerfScout - An Adaptive Workload Generator in Software Performance Testing]] - 所属: [[Nankai University]] - 共同研究者: [[Shenglin Zhang]] / [[Yongxin Zhao]] / [[Shiyu Ma]] / [[Tong Xiao]] / [[Dan Pei]] / [[Yu Luo]] / [[Sibo Xia]] / [[Yuan Yuan]] / [[Tongqing Zhou]] / [[Chenyu Zhao]] / [[Wenwei Gu]] / [[Weilin Jin]] / [[Ruowei Fu]] / [[Junhua Kuang]] / [[Yimeng Zhang]] / [[Jintao Feng]] / [[Jingyu Wang]] / [[Yuxin Sun]] / [[Xidao Wen]] / [[Jiacheng Zhang]] / [[Guohua Liu]] / [[Shiqi Chen]] / [[Minghua Ma]] / [[Yutong Chen]] / [[Tianyu Cui]] / [[Yicheng Sui]] / [[Zeyu Che]] - 関連プロダクト: [[LogInsight]] / [[OptProphet]] / [[SCELM]] / [[OpsAgent]] / [[OpsMem]] / [[SuperAgg]] / [[UniDiag]] / [[AgentTether]] / [[Build-bench]] / [[OncallX]] / CoTriage / [[LagRCA]] / [[OScope]] / [[PerfScout]] / [[RefinedEdge]] / LogSage - 概念: [[ログベース障害診断]] / [[マルチモーダル障害診断]] / [[時系列知識グラフ]] / [[エージェント修復]] / [[クロスISAマイグレーション]] / [[自動ビルド修復]] / [[オンコール自動化]] / [[知識蒸留]] / [[インシデントトリアージ]] / [[TSG自動化]] / [[定常性モデル]] / [[異常検知]] / [[モデル圧縮]] / [[Edge-cloud Collaboration]] / [[グラフベースRCA]]