# Tsinghua University
中国の研究大学。[[Minder]] 論文([[@2025__NSDI__Minder - Faulty Machine Detection for Large-scale Distributed Model Training]])の筆頭著者 [[Yangtao Deng]] と Xingjian Zhang の所属で、本研究は [[ByteDance]] との共同研究として行われた。(Source: [[@2025__NSDI__Minder - Faulty Machine Detection for Large-scale Distributed Model Training]])
[[Fastsocket]] を扱う ASPLOS '16 論文では、[[Yu Chen (Tsinghua)|Yu Chen]]、[[Junjie Mao]]、[[Jiaquan He]]、[[Wei Xu]]、[[Yuanchun Shi]] の所属として記載されている。論文は Tsinghua University と Sina Corporation・ZHIHU Corporation の協力による、BSD Socket 互換のスケーラブルなカーネル TCP スタック設計を報告する。(Source: [[@2016__ASPLOS__Scalable Kernel TCP Design and Implementation for Short-Lived Connections]])
- [[Stratus]] 一次論文([[@2025__NeurIPS2025__STRATUS - A Multi-agent System for Autonomous Reliability Engineering of Modern Clouds]])でも所属 `¶`(IIIS, Tsinghua University)として、共著代表の一人 Jiaqi Pan の兼務所属に挙がる(主所属は [[University of Illinois Urbana-Champaign]])。(Source: [[@2025__NeurIPS2025__STRATUS - A Multi-agent System for Autonomous Reliability Engineering of Modern Clouds]])
- [[A Survey of AIOps in the Era of Large Language Models]] の共著者 Aiwei Liu(
[email protected])の所属でもある(主所属は [[Peking University]] グループとの共同)。(Source: [[A Survey of AIOps in the Era of Large Language Models]])
- [[LagRCA]] 論文([[@2026__FSE Companion__Bridging the Delay - Lag-Aware Spatio-Temporal Causal Inference for Microservice Root Cause Analysis]], FSE Companion '26)の所属の一つ(& BNRist)。[[Dan Pei]] が [[Nankai University]] の [[Shenglin Zhang]]・[[Yongqian Sun]] らと共同で、マイクロサービス障害の可変時間ラグを明示的にモデル化する遅延認識時空間因果推論フレームワークを提案した。(Source: [[@2026__FSE Companion__Bridging the Delay - Lag-Aware Spatio-Temporal Causal Inference for Microservice Root Cause Analysis]])
- [[OpenRCA]] 論文([[@2025__ICLR__OpenRCA - Can Large Language Models Locate the Root Cause of Software Failures]], ICLR 2025)の所属の一つ。AIOps を長く牽引する [[Dan Pei]](
[email protected])が共著者として参加した。(Source: [[@2025__ICLR__OpenRCA - Can Large Language Models Locate the Root Cause of Software Failures]])
- [[Hawkeye]] 論文([[@2025__SIGCOMM__Hawkeye - Diagnosing RDMA Network Performance Anomalies with PFC Provenance]], SIGCOMM 2025)の主たる所属。筆頭著者 [[Shicheng Wang]] をはじめ [[Mingwei Xu]]・[[Jiahai Yang]]・Xiao Li・Zhiliang Wang・Xingang Shi が在籍し、[[Beihang University]]・[[Infrawaves]] との共同で RDMA 性能異常の PFC プロベナンス診断に取り組む。(Source: [[@2025__SIGCOMM__Hawkeye - Diagnosing RDMA Network Performance Anomalies with PFC Provenance]])
- [[SkeletonHunter]] 論文([[@2025__SIGCOMM__SkeletonHunter - Diagnosing and Localizing Network Failures in Containerized Large Model Training]], SIGCOMM 2025)の所属の一つ。筆頭著者 Wei Liu(Alibaba Cloud との兼務)・責任著者 Zhenhua Li・Yunhao Liu が在籍し、Alibaba Cloud・[[University of Illinois Urbana-Champaign]] と共同でコンテナネットワーク障害診断に取り組む。(Source: [[@2025__SIGCOMM__SkeletonHunter - Diagnosing and Localizing Network Failures in Containerized Large Model Training]])
- [[D-Bot]] 論文([[@2024__PVLDB__D-Bot - Database Diagnosis System using Large Language Models]], PVLDB 2024)の主たる所属。[[Xuanhe Zhou]]・[[Guoliang Li]] らの Database Group が LLM ベースのデータベース異常診断システムを開発した([[DB-GPT]] リポジトリ)。(Source: [[@2024__PVLDB__D-Bot - Database Diagnosis System using Large Language Models]])
- [[EcoTune]] 論文([[@2025__SIGMOD__Rethinking The Compaction Policies in LSM-trees]], SIGMOD 2025)の全著者 [[Hengrui Wang]]・[[Jiansheng Qiu]]・[[Fangzhou Yuan]]・[[Huanchen Zhang]] の所属。LSM ツリーのコンパクション方針を平均クエリスループット最適化として再定式化した。(Source: [[@2025__SIGMOD__Rethinking The Compaction Policies in LSM-trees]])
- [[@2026__arXiv__Position - The Inevitable End of One-Architecture-Fits-All-Domains in Time Series Forecasting]] の著者 4 名中 3 名([[Qinwei Ma]]・[[Jingzhe Shi]]・[[Zaiwen Yang]])の所属。時系列予測における汎ドメインアーキテクチャの限界を論じたポジションペーパー。(Source: [[@2026__arXiv__Position - The Inevitable End of One-Architecture-Fits-All-Domains in Time Series Forecasting]])
- [[@2022__ACL__GLM - General Language Model Pretraining with Autoregressive Blank Infilling|GLM]] 論文([[@2022__ACL__GLM - General Language Model Pretraining with Autoregressive Blank Infilling]], ACL 2022)の主たる所属。筆頭著者 [[Zhengxiao Du]]・共著者 [[Xiao Liu]]・[[Ming Ding]]・[[Jiezhong Qiu]] が在籍し、責任著者 [[Jie Tang]]・[[Zhilin Yang]] の指導下で自己回帰空白埋めによる汎用言語モデルフレームワークを開発した。本成果は後の GLM-130B・ChatGLM・GLM-4 ファミリーへと発展する。(Source: [[@2022__ACL__GLM - General Language Model Pretraining with Autoregressive Blank Infilling]])
- [[ChatTS]] 論文([[@2025__VLDB__ChatTS - Aligning Time Series with LLMs via Synthetic Data for Enhanced Understanding and Reasoning]], PVLDB Vol. 18, 2025)の主たる所属(BNRist と並記)。筆頭著者 [[Zhe Xie]]・共著者 [[Longlong Xu]]・corresponding author [[Dan Pei]] が在籍し、ByteDance・BizSeer との共同で時系列マルチモーダル LLM の初実装を構築。NetManAIOps グループの「AIOps × 時系列 + LLM」研究の到達点の一つ。(Source: [[@2025__VLDB__ChatTS - Aligning Time Series with LLMs via Synthetic Data for Enhanced Understanding and Reasoning]])
- **Mooncake** 論文([[@2024__arXiv__Mooncake - A KVCache-centric Disaggregated Architecture for LLM Serving]], arXiv 2024)の共同機関。MadSys グループの [[Mingxing Zhang]]・[[Yongwei Wu]]・[[Weimin Zheng]] が [[Moonshot AI]] と共同で、KVCache 中心の分散 LLM サービングアーキテクチャを開発。Prefill/Decode 分離・CPU/DRAM/SSD 分散 KVCache プール・Chunked Pipeline Parallelism などを提案し、実ワークロードで vLLM 比 75% 多いリクエスト処理を実証。(Source: [[@2024__arXiv__Mooncake - A KVCache-centric Disaggregated Architecture for LLM Serving]])
- **KVShare** 論文([[@2025__arXiv__KVShare - An LLM Service System with Efficient and Effective Multi-Tenant KV Cache Reuse]], arXiv 2025)の共同機関。Mingzhe Huang・Weijun Wang・Yuanchun Li・Yunxin Liu が在籍し、[[Central South University]] と共同でマルチテナント KV キャッシュ共有フレームワーク [[KVShare]] を開発。DHD アルゴリズムと cache-aware スケジューラにより TTFT 最大 9.39 倍短縮・SOTA 比 20.38% 精度改善を達成した。(Source: [[@2025__arXiv__KVShare - An LLM Service System with Efficient and Effective Multi-Tenant KV Cache Reuse]])
[[Dan Pei]](博士課程学生・指導教員の輩出元)が Build-bench 論文([[@2026__TOSEM__Can Language Models Go Beyond Coding - Assessing the Capability of Language Models to Build Real-World Systems]], nkcs.iops.ai 2026-05)に共著者として参加。[[Nankai University]]・[[Peking University]]・[[Microsoft]] とのクロス ISA(x86_64/aarch64)ビルド修復ベンチマーク開発に寄与した。(Source: [[@2026__TOSEM__Can Language Models Go Beyond Coding - Assessing the Capability of Language Models to Build Real-World Systems]])
[[Dan Pei]] が OScope 論文([[@2026__ICSE-SEIP__When LLMs Listen to Experts - Accurate Failure Diagnosis in Operating Systems]], ICSE-SEIP '26)に共著者として参加。[[Nankai University]]([[Yongxin Zhao]] 筆頭・[[Shenglin Zhang]] 責任著者)・[[Alibaba Group]] との共同で、OS 障害診断向け LLM フレームワーク [[OScope]] の開発に寄与した。(Source: [[@2026__ICSE-SEIP__When LLMs Listen to Experts - Accurate Failure Diagnosis in Operating Systems]])
[[Dan Pei]] が [[PerfScout]] 論文([[@2026__ICSE-SEIP__PerfScout - An Adaptive Workload Generator in Software Performance Testing]], ICSE-SEIP '26)に共著者として参加。[[Nankai University]]([[Yongqian Sun]]・[[Shenglin Zhang]] 責任著者ほか)・[[BizSeer]]([[Xidao Wen]])・[[Huawei Cloud]](成都)との共同で、SPOT(極値理論)・ADF/KPSS(局所定常性判定)・PPO(強化学習)を統合した性能テスト向け適応的ワークロード生成フレームワーク [[PerfScout]] の開発に寄与した。Huawei Cloud の CodeArts PerfTest に 9 か月間本番デプロイされ、ブレークポイント特定精度 82% 超を達成。RCA・異常検知が中心だった NetManAIOps グループの研究射程が性能テスト自動化にも広がったことを示す。(Source: [[@2026__ICSE-SEIP__PerfScout - An Adaptive Workload Generator in Software Performance Testing]])
[[Dan Pei]]・[[Qingyi Guo]] が OpsMem 論文([[@2026__arXiv__OpsMem - Dual-Memory Reasoning with Cross-Memory Resonance for Failure Diagnosis]], arXiv 2026-07)に共著者として参加。[[Nankai University]]([[Yongqian Sun]] 筆頭・[[Shenglin Zhang]] 責任著者ほか)・[[Huawei Technologies]] との共同で、短期記憶と長期記憶を cross-memory resonance で結合する失敗診断向けデュアルメモリフレームワーク [[OpsMem]] の開発に寄与した。(Source: [[@2026__arXiv__OpsMem - Dual-Memory Reasoning with Cross-Memory Resonance for Failure Diagnosis]])
[[Yihe Wang]](Department of Computer Science and Technology)が LLM ハルシネーション引用の大規模監査研究([[@2026__arXiv__LLM hallucinations in the wild]]、arXiv 2026-05-08)に共著者として参加。[[Cornell University]]・[[University of California, Berkeley]] のグループと共同で arXiv・bioRxiv・SSRN・PubMed Central 4 コーパスを横断監査し、2025 年単年で 146,932 件のハルシネーション引用を推定した。従来の AIOps/システム系ソースとは異なる、科学計量学(science of science)分野からの新規ソースである。(Source: [[@2026__arXiv__LLM hallucinations in the wild]])
[[Dan Pei]] が HeaRank 論文([[@2026__arXiv__Don't Predict, Prioritize - Rethinking GPU Reliability Assessment]]、KDD '26 V.2)の Computer Science Department 所属著者として参加。CNIC/CAS の [[Changhua Pei]]・[[Gaogang Xie]] らおよび [[StepFun]] の [[Yibo Zhu]] との共同で、GPU 障害の時系列予測が本質的に困難であることを実証し、Learning-to-Rank によるホストリスクランキングモデル HeaRank を提案した。マイクロサービス RCA・時系列異常検知が中心だった NetManAIOps グループの研究射程に、GPU ハードウェア信頼性という新ドメインが加わった。(Source: [[@2026__arXiv__Don't Predict, Prioritize - Rethinking GPU Reliability Assessment]])
[[Nengwen Zhao]]・[[Xiaohui Nie]]・[[Dan Pei]](BNRist 併記)が eWarn 論文([[@2020__ESEC-FSE__Real-Time Incident Prediction for Online Service Systems]], ESEC/FSE '20)に参加。[[Tianjin University]] の [[Junjie Chen]](対応著者)・[[BizSeer]]([[Zhou Wang]]・[[Wenchi Zhang]]・[[Kaixin Sui]])・[[China EverBright Bank]] との産学連携で、アラートデータのみからリアルタイムにインシデントを予測する手法 eWarn を提案した。LDA テキスト特徴+統計特徴+multi-instance learning でノイズアラートを抑制し、大手商業銀行 11 実サービスシステムで平均 F1 0.82(AirAlert 比 +0.31)を達成、2 行への実運用適用も報告した。NetManAIOps グループのアラートストーム研究(2020)に続く、alert-based な先回り型障害対処研究の系譜。(Source: [[@2020__ESEC-FSE__Real-Time Incident Prediction for Online Service Systems]])
[[Peng Cui]] が [[DAG-FM]] 論文([[@2026__arXiv__DAG-FM - A Foundation Model for Causal Discovery under Heterogeneous Causal Mechanisms]]、arXiv 2026-07)に共著者として参加。[[Zhejiang University]] の [[Kun Kuang]](責任著者)・Yikang Chen・Zhengkang Guan・Haoyuan Qian・Yi Yang との共同で、異種因果メカニズム下でも識別可能な DAG を保証する因果発見基盤モデルを提案した。本頁の他エントリの多くを占める NetManAIOps グループ([[Dan Pei]] 系、RCA・異常検知中心)とは異なる、因果機械学習の研究系譜として新規に接続された。(Source: [[@2026__arXiv__DAG-FM - A Foundation Model for Causal Discovery under Heterogeneous Causal Mechanisms]])
AnoFusion 論文([[@2023__KDD__Robust Multimodal Failure Detection for Microservice Systems]]、KDD '23)の共著者所属。[[Dan Pei]] が [[Nankai University]]・[[Microsoft]] と共同で、metric/log/trace のマルチモーダル相関を GTN+GAT で学習する教師なしインスタンス障害検知手法 [[AnoFusion]] を提案した。(Source: [[@2023__KDD__Robust Multimodal Failure Detection for Microservice Systems]])
[[Chang Liu]]・[[Long Wang]]([[Zhongguancun Laboratory]]兼務)が eBPF マップ性能ベンチマーク論文([[@2024__eBPF'24__Understanding Performance of eBPF Maps]]、eBPF '24)の著者として参加。[[Kyungpook National University]]の[[Byungchul Tak]]との共同で、eBPFマップのアクセスオーバーヘッドを体系的にベンチマークし、メモリフットプリントとキャッシュホット性が主要因であることを示した。(Source: [[@2024__eBPF'24__Understanding Performance of eBPF Maps]])
[[Leyi Pan]] が [[Bifrost]] 論文([[@2026__arXiv__Bifrost - Empowering Pretrained Language Model with Fallibility Representation for Log-Based Fault Diagnosis]], ASE '26 採録)に共著者として参加。[[Peking University]]([[Minghua He]] 筆頭・[[Tong Jia]]/[[Ying Li]] corresponding)・[[Alibaba Group]] との共同で、ログの多階層構造(実行フロー・イベント・コンポーネント)を fallibility representation として学習する対照学習手法を提案した。既存の NetManAIOps 系(Dan Pei グループ)の RCA・時系列異常検知研究とは独立に、ログ表現学習という新たな観点から PKU グループの研究に参加した点が特徴。(Source: [[@2026__arXiv__Bifrost - Empowering Pretrained Language Model with Fallibility Representation for Log-Based Fault Diagnosis]])
Zili Meng・Lianjin Ye・Jingyu Xiao・Jilong Wang(BNRist・Peng Cheng Laboratory兼務)・Heng Yuが、[[Tencent]]・[[UC Santa Barbara]]と共同で光バックボーンネットワーク向けテレメトリシステムOpTel([[@2022__NSDI__Detecting Ephemeral Optical Events with OpTel]], NSDI 2022)を発表。SNMPベースの既存テレメトリの限界を、ベンダー非依存の標準化デバイスモデルとpush型テレメトリパイプラインで解消し、Tencentの光バックボーンで6か月間の本番運用実績を報告する。既存のマイクロサービス・LLM訓練系AIOps研究とは異なる、物理層(光ネットワーク)のテレメトリ研究として本頁に新規に接続される。(Source: [[@2022__NSDI__Detecting Ephemeral Optical Events with OpTel]])
[[Dan Pei]]・[[Zeyan Li]]・[[Nengwen Zhao]]・[[Xidao Wen]] が [[@2022__ESEC FSE__Constructing Large-Scale Real-World Benchmark Datasets for AIOps]](ESEC/FSE 2022 Industry Track)の主たる所属(¹)として参加。[[Nankai University]]([[Shenglin Zhang]]・[[Yongqian Sun]])・[[Sun Yat-sen University]]([[Pengfei Chen]])・[[Microsoft Research]]([[Minghua Ma]])との共同で、KPI 異常検知・多次元根本原因箇所特定・障害発見/診断の 3 公開データセットと年次 AIOps アルゴリズムコンペティションを紹介した。(Source: [[@2022__ESEC FSE__Constructing Large-Scale Real-World Benchmark Datasets for AIOps]])
[[Dan Pei]] が RouterOPS 論文([[@2026__SIGCOMM Posters and Demos__Predict Boldly, Recover Cautiously : Fast On-Router Route Anomaly Prediction and Recovery]]、SIGCOMM Posters and Demos '26)に共著者として参加。CNIC/CAS の [[Hang Cui]](筆頭)・[[Cenjie Hu]]・[[Zexin Wang]]・[[Jingjing Li]]・[[Juncheng Hu]]・[[Changhua Pei]]・[[Gaogang Xie]] との共同で、ルーターが自己/近隣の異常を検知し安全ガード付きでローカル復旧するルーターネイティブ軽量 AIOps フレームワークを提案した。NetManAIOps グループの研究射程が、マイクロサービス RCA・時系列異常検知・GPU 信頼性・性能テストに続き、ネットワークルーティング層の障害予測・復旧にも及んだ。(Source: [[@2026__SIGCOMM Posters and Demos__Predict Boldly, Recover Cautiously : Fast On-Router Route Anomaly Prediction and Recovery]])
[[Dan Pei]]・[[Yuhe Liu]]・[[Longlong Xu]] が Eagle 論文([[@2026__FSE Companion__Eagle - Leveraging Operations Documents for Comprehensive Benchmark Question Generation]]、FSE Companion '26)の Tsinghua/BNRist 側著者として参加。CNIC/CAS([[Changhua Pei]]・[[Hang Wang]])・[[Huawei Technologies]]・[[China Academy of Information and Communications Technology]] との共同で、Ops ドキュメントから運用中心のベンチマーク QA を自動生成するフレームワーク [[Eagle (OpsLLMベンチマーク)]] を開発し、Huawei 社内に6ヶ月間デプロイして4,845件の QA ペアを合成した。NetManAIOps グループの研究射程が RCA・異常検知の「診断対象」から OpsLLM 自体を評価するベンチマーク基盤の構築へ広がったことを示す一本。(Source: [[@2026__FSE Companion__Eagle - Leveraging Operations Documents for Comprehensive Benchmark Question Generation]])
[[Yong Liu]]・[[Guo Qin]]・[[Xiangdong Huang]]・[[Jianmin Wang]]・[[Mingsheng Long]](責任著者)が Timer-XL 論文([[@2025__ICLR__Timer-XL - Long-Context Transformers for Unified Time Series Forecasting]]、ICLR 2025)の全著者として参加。School of Software, BNRist が主たる所属で、単変量次トークン予測を多変量へ一般化した TimeAttention により、単変量・多変量・共変量付き予測を統一する decoder-only Transformer を提案した。Mingsheng Long 研究室が継続する時系列 Transformer 研究(Autoformer・TimesNet・iTransformer・Timer)の系譜に連なる。既存の Dan Pei 系(NetManAIOps、AIOps・RCA 中心)とは異なる、[[多変量時系列予測]]・[[時系列基盤モデル]]を専門とする研究グループとして本頁に新規に接続される。(Source: [[@2025__ICLR__Timer-XL - Long-Context Transformers for Unified Time Series Forecasting]])
[[Zhipu AI|Z.ai]]・[[Harnets.AI]] と共同で、ネットワークトポロジ [[ZCube]] を GLM-5.1 コーディング推論の本番クラスタへ展開した技術記事([[@2026__X__Next-generation LLM Inference Network - How ZCube Alleviates Network Bottlenecks]]、2026-05-20)の共同開発機関として参加。ZCube は ACM SIGCOMM 2025 で発表された、Spine 層を撤廃した完全フラット化トポロジで、Prefill-Decode 分離推論の KV Cache 転送非対称性が引き起こすトポロジ誘発輻輳を解消する。既存の [[Dan Pei]] 系 NetManAIOps グループ(マイクロサービス RCA・時系列異常検知中心)や THUML グループ(時系列基盤モデル)とは異なる、データセンターネットワークアーキテクチャ研究の系譜として本頁に新規に接続される。(Source: [[@2026__X__Next-generation LLM Inference Network - How ZCube Alleviates Network Bottlenecks]])
[[Mingsheng Long]] 率いる THUML グループ(School of Software, BNRist)が、[[Sundial]] 論文([[@2025__ICML__Sundial - A Family of Highly Capable Time Series Foundation Models]], ICML 2025)で flow-matching ベースの生成的[[時系列基盤モデル]]を発表。筆頭著者 [[Yong Liu]]・[[Guo Qin]](同等貢献)・共著者 Zhiyuan Shi・Zhi Chen・Caiyin Yang・[[Xiangdong Huang]]・[[Jianmin Wang]] が参加し、1 兆時系列点規模の TimeBench で TSLib・GIFT-Eval・FEV leaderboard のゼロショット SOTA を達成した。同グループは [[@2025__ICLR__Timer-XL - Long-Context Transformers for Unified Time Series Forecasting|Timer-XL]]・iTransformer・Koopa・Non-stationary Transformers など時系列 Transformer 研究を継続的に発表しており、NetManAIOps グループ([[Dan Pei]] 系、RCA・異常検知中心)とは異なる、時系列基盤モデル・アーキテクチャ研究の系譜として本頁に新規に接続される。(Source: [[@2025__ICML__Sundial - A Family of Highly Capable Time Series Foundation Models]])
[[Dan Pei]] がマイクロサービス障害診断包括サーベイ([[Failure Diagnosis in Microservice Systems]], arXiv 2024)に Tsinghua University 側著者として参加。[[Shenglin Zhang]]([[Nankai University]])を筆頭著者とするチームに [[Minghua Ma]]([[Microsoft]])・[[Yongqian Sun]] とともに参加し、2003 年から現在までの 98 論文を対象に根本原因箇所特定(RCL)と障害分類(FC)を区別する問題定式化・マルチモーダルデータ taxonomy・公開データセット/ツールキット/評価指標の体系化を行った。(Source: [[@2024__arXiv__Failure Diagnosis in Microservice Systems - A Comprehensive Survey and Analysis - Chapter 1 Introduction]], [[@2024__arXiv__Failure Diagnosis in Microservice Systems - A Comprehensive Survey and Analysis - Chapter 3 Terminologies]])
## 関連
- ソース: [[@2025__ICML__Sundial - A Family of Highly Capable Time Series Foundation Models]] / [[@2026__SIGCOMM Posters and Demos__Predict Boldly, Recover Cautiously : Fast On-Router Route Anomaly Prediction and Recovery]] / [[@2026__X__Next-generation LLM Inference Network - How ZCube Alleviates Network Bottlenecks]] / [[@2026__ISSRE__CoLMAD : Cost-Efficient LLM-Assisted Time Series Anomaly Detection for Industrial Monitoring]] / [[@2026__SIGCOMM__CubeTrace Microscopic Network Tracing for Heterogeneous Cloud Gateways]] / [[@2026__WWW2026__ViTs - Teaching Machines to See Time Series Anomalies Like Human Experts]] / [[@2026__ICLR__AutoDA-Timeseries - Automated Data Augmentation for Time Series]] / [[@2025__NSDI__Mitigating Scalability Walls of RDMA-based Container Networks]] / [[@2026__SIGCOMM__Pegasus - A Data Center Network for Bare-Metal AI Cloud]] / [[@2026__SIGCOMM__Networked Agent Memory and Causality Representation - Experiences towards Interpretable Cloud-Scale Root-Causing]]
- ソース: [[@2026__ISSRE__ChainCraft - Bridging Causal Discovery and LLM Reasoning for Failure Prediction in Microservices]] / [[@2026__FSE Companion__Eagle - Leveraging Operations Documents for Comprehensive Benchmark Question Generation]] / [[@2023__KDD__Robust Multimodal Failure Detection for Microservice Systems]] / [[@2026__arXiv__Don't Predict, Prioritize - Rethinking GPU Reliability Assessment]] / [[@2026__TOSEM__Can Language Models Go Beyond Coding - Assessing the Capability of Language Models to Build Real-World Systems]] / [[@2026__arXiv__OpsMem - Dual-Memory Reasoning with Cross-Memory Resonance for Failure Diagnosis]] / [[@2026__ICSE-SEIP__When LLMs Listen to Experts - Accurate Failure Diagnosis in Operating Systems]] / [[@2025__NSDI__Minder - Faulty Machine Detection for Large-scale Distributed Model Training]] / [[@2025__NeurIPS2025__STRATUS - A Multi-agent System for Autonomous Reliability Engineering of Modern Clouds]] / [[A Survey of AIOps in the Era of Large Language Models]] / [[@2025__ICLR__OpenRCA - Can Large Language Models Locate the Root Cause of Software Failures]] / [[@2025__SIGCOMM__Hawkeye - Diagnosing RDMA Network Performance Anomalies with PFC Provenance]] / [[@2025__SIGCOMM__SkeletonHunter - Diagnosing and Localizing Network Failures in Containerized Large Model Training]] / [[@2024__PVLDB__D-Bot - Database Diagnosis System using Large Language Models]] / [[@2025__SIGMOD__Rethinking The Compaction Policies in LSM-trees]] / [[@2026__arXiv__Position - The Inevitable End of One-Architecture-Fits-All-Domains in Time Series Forecasting]] / [[@2022__ACL__GLM - General Language Model Pretraining with Autoregressive Blank Infilling]] / [[@2025__arXiv__KVShare - An LLM Service System with Efficient and Effective Multi-Tenant KV Cache Reuse]] / [[@2026__ICSE-SEIP__PerfScout - An Adaptive Workload Generator in Software Performance Testing]] / [[@2026__arXiv__Bifrost - Empowering Pretrained Language Model with Fallibility Representation for Log-Based Fault Diagnosis]] / [[@2025__ICLR__Timer-XL - Long-Context Transformers for Unified Time Series Forecasting]]
- ノブチューニングサーベイ([[@2023__TKDE__Automatic Database Knob Tuning - A Survey]], IEEE TKDE 2023)の全著者 [[Xinyang Zhao]]・[[Xuanhe Zhou]]・[[Guoliang Li]] の所属。ノブチューニングのパイプラインを4段階に分解し、16手法を体系的に比較した初の包括的サーベイ。(Source: [[@2023__TKDE__Automatic Database Knob Tuning - A Survey]])
ChainCraft 論文([[@2026__ISSRE__ChainCraft - Bridging Causal Discovery and LLM Reasoning for Failure Prediction in Microservices]], ISSRE 2026)は [[Dan Pei]] が Tsinghua University から参加し、[[Nankai University]]([[Yongxin Zhao]]・[[Yongqian Sun]] ら)・[[Alibaba Group]]([[Boxuan Zhao]]・[[Li Shi]]・[[Wei Li]]・[[Liping Zhang]])と共同開発した。メトリクス駆動の因果発見(PCMCI)と LLM 推論を橋渡しする障害予測フレームワークで、Alibaba 本番データ(F1=0.875)と3か月間の産業デプロイで有効性を実証した。(Source: [[@2026__ISSRE__ChainCraft - Bridging Causal Discovery and LLM Reasoning for Failure Prediction in Microservices]])
- エンティティ: [[Yangtao Deng]] / [[ByteDance]] / [[University of Illinois Urbana-Champaign]] / [[IBM Research]] / [[Peking University]] / [[Alibaba Group]] / [[SkeletonHunter]] / [[Xuanhe Zhou]] / [[Guoliang Li]] / [[DB-GPT]] / [[Hengrui Wang]] / [[Huanchen Zhang]] / [[EcoTune]] / [[Qinwei Ma]] / [[Jingzhe Shi]] / [[Zaiwen Yang]] / [[Xinyang Zhao]] / [[Nankai University]] / [[BizSeer]] / [[Huawei Cloud]] / [[PerfScout]] / [[Yihe Wang]] / [[Cornell University]] / [[University of California, Berkeley]] / [[Yong Liu]] / [[Guo Qin]] / [[Xiangdong Huang]] / [[Jianmin Wang]] / [[Mingsheng Long]] / [[Zhipu AI]] / [[Harnets.AI]]
- 概念: [[定常性モデル]] / [[ZCube]]
- [[@2026__ICLR__AutoDA-Timeseries - Automated Data Augmentation for Time Series]](ICLR 2026)。[[Dan Pei]] グループ(NetManAIOps)による時系列向け自動データ拡張フレームワークの研究。(Source: [[@2026__ICLR__AutoDA-Timeseries - Automated Data Augmentation for Time Series]])
## 出典
- [[@2026__SIGCOMM__CubeTrace Microscopic Network Tracing for Heterogeneous Cloud Gateways]](Mingwei Xuの所属(SIGCOMM 2026))
- [[@2026__SIGCOMM__Pegasus - A Data Center Network for Bare-Metal AI Cloud]](筆頭著者 Xianneng Zou の所属。ベアメタルAIクラウド向けデータセンターネットワークPegasusの共著。)