# Tsinghua University 中国の研究大学。[[Minder]] 論文([[@2025__NSDI__Minder - Faulty Machine Detection for Large-scale Distributed Model Training]])の筆頭著者 [[Yangtao Deng]] と Xingjian Zhang の所属で、本研究は [[ByteDance]] との共同研究として行われた。(Source: [[@2025__NSDI__Minder - Faulty Machine Detection for Large-scale Distributed Model Training]]) - [[Stratus]] 一次論文([[@2025__NeurIPS2025__STRATUS - A Multi-agent System for Autonomous Reliability Engineering of Modern Clouds]])でも所属 `¶`(IIIS, Tsinghua University)として、共著代表の一人 Jiaqi Pan の兼務所属に挙がる(主所属は [[University of Illinois Urbana-Champaign]])。(Source: [[@2025__NeurIPS2025__STRATUS - A Multi-agent System for Autonomous Reliability Engineering of Modern Clouds]]) - [[@2025__CSUR__A Survey of AIOps in the Era of Large Language Models]] の共著者 Aiwei Liu([email protected])の所属でもある(主所属は [[Peking University]] グループとの共同)。(Source: [[@2025__CSUR__A Survey of AIOps in the Era of Large Language Models]]) - [[LagRCA]] 論文([[@2026__FSE Companion__Bridging the Delay - Lag-Aware Spatio-Temporal Causal Inference for Microservice Root Cause Analysis]], FSE Companion '26)の所属の一つ(& BNRist)。[[Dan Pei]] が [[Nankai University]] の [[Shenglin Zhang]]・[[Yongqian Sun]] らと共同で、マイクロサービス障害の可変時間ラグを明示的にモデル化する遅延認識時空間因果推論フレームワークを提案した。(Source: [[@2026__FSE Companion__Bridging the Delay - Lag-Aware Spatio-Temporal Causal Inference for Microservice Root Cause Analysis]]) - [[OpenRCA]] 論文([[@2025__ICLR__OpenRCA - Can Large Language Models Locate the Root Cause of Software Failures]], ICLR 2025)の所属の一つ。AIOps を長く牽引する [[Dan Pei]]([email protected])が共著者として参加した。(Source: [[@2025__ICLR__OpenRCA - Can Large Language Models Locate the Root Cause of Software Failures]]) - [[Hawkeye]] 論文([[@2025__SIGCOMM__Hawkeye - Diagnosing RDMA Network Performance Anomalies with PFC Provenance]], SIGCOMM 2025)の主たる所属。筆頭著者 [[Shicheng Wang]] をはじめ [[Mingwei Xu]]・[[Jiahai Yang]]・Xiao Li・Zhiliang Wang・Xingang Shi が在籍し、[[Beihang University]]・[[Infrawaves]] との共同で RDMA 性能異常の PFC プロベナンス診断に取り組む。(Source: [[@2025__SIGCOMM__Hawkeye - Diagnosing RDMA Network Performance Anomalies with PFC Provenance]]) - [[SkeletonHunter]] 論文([[@2025__SIGCOMM__SkeletonHunter - Diagnosing and Localizing Network Failures in Containerized Large Model Training]], SIGCOMM 2025)の所属の一つ。筆頭著者 Wei Liu(Alibaba Cloud との兼務)・責任著者 Zhenhua Li・Yunhao Liu が在籍し、Alibaba Cloud・[[University of Illinois Urbana-Champaign]] と共同でコンテナネットワーク障害診断に取り組む。(Source: [[@2025__SIGCOMM__SkeletonHunter - Diagnosing and Localizing Network Failures in Containerized Large Model Training]]) - [[D-Bot]] 論文([[@2024__PVLDB__D-Bot - Database Diagnosis System using Large Language Models]], PVLDB 2024)の主たる所属。[[Xuanhe Zhou]]・[[Guoliang Li]] らの Database Group が LLM ベースのデータベース異常診断システムを開発した([[DB-GPT]] リポジトリ)。(Source: [[@2024__PVLDB__D-Bot - Database Diagnosis System using Large Language Models]]) - [[EcoTune]] 論文([[@2025__SIGMOD__Rethinking The Compaction Policies in LSM-trees]], SIGMOD 2025)の全著者 [[Hengrui Wang]]・[[Jiansheng Qiu]]・[[Fangzhou Yuan]]・[[Huanchen Zhang]] の所属。LSM ツリーのコンパクション方針を平均クエリスループット最適化として再定式化した。(Source: [[@2025__SIGMOD__Rethinking The Compaction Policies in LSM-trees]]) - [[@2026__arXiv__Position - The Inevitable End of One-Architecture-Fits-All-Domains in Time Series Forecasting]] の著者 4 名中 3 名([[Qinwei Ma]]・[[Jingzhe Shi]]・[[Zaiwen Yang]])の所属。時系列予測における汎ドメインアーキテクチャの限界を論じたポジションペーパー。(Source: [[@2026__arXiv__Position - The Inevitable End of One-Architecture-Fits-All-Domains in Time Series Forecasting]]) - [[@2022__ACL__GLM - General Language Model Pretraining with Autoregressive Blank Infilling|GLM]] 論文([[@2022__ACL__GLM - General Language Model Pretraining with Autoregressive Blank Infilling]], ACL 2022)の主たる所属。筆頭著者 [[Zhengxiao Du]]・共著者 [[Xiao Liu]]・[[Ming Ding]]・[[Jiezhong Qiu]] が在籍し、責任著者 [[Jie Tang]]・[[Zhilin Yang]] の指導下で自己回帰空白埋めによる汎用言語モデルフレームワークを開発した。本成果は後の GLM-130B・ChatGLM・GLM-4 ファミリーへと発展する。(Source: [[@2022__ACL__GLM - General Language Model Pretraining with Autoregressive Blank Infilling]]) - [[ChatTS]] 論文([[@2025__VLDB__ChatTS - Aligning Time Series with LLMs via Synthetic Data for Enhanced Understanding and Reasoning]], PVLDB Vol. 18, 2025)の主たる所属(BNRist と並記)。筆頭著者 [[Zhe Xie]]・共著者 [[Longlong Xu]]・corresponding author [[Dan Pei]] が在籍し、ByteDance・BizSeer との共同で時系列マルチモーダル LLM の初実装を構築。NetManAIOps グループの「AIOps × 時系列 + LLM」研究の到達点の一つ。(Source: [[@2025__VLDB__ChatTS - Aligning Time Series with LLMs via Synthetic Data for Enhanced Understanding and Reasoning]]) - **Mooncake** 論文([[@2024__arXiv__Mooncake - A KVCache-centric Disaggregated Architecture for LLM Serving]], arXiv 2024)の共同機関。MadSys グループの [[Mingxing Zhang]]・[[Yongwei Wu]]・[[Weimin Zheng]] が [[Moonshot AI]] と共同で、KVCache 中心の分散 LLM サービングアーキテクチャを開発。Prefill/Decode 分離・CPU/DRAM/SSD 分散 KVCache プール・Chunked Pipeline Parallelism などを提案し、実ワークロードで vLLM 比 75% 多いリクエスト処理を実証。(Source: [[@2024__arXiv__Mooncake - A KVCache-centric Disaggregated Architecture for LLM Serving]]) - **KVShare** 論文([[@2025__arXiv__KVShare - An LLM Service System with Efficient and Effective Multi-Tenant KV Cache Reuse]], arXiv 2025)の共同機関。Mingzhe Huang・Weijun Wang・Yuanchun Li・Yunxin Liu が在籍し、[[Central South University]] と共同でマルチテナント KV キャッシュ共有フレームワーク [[KVShare]] を開発。DHD アルゴリズムと cache-aware スケジューラにより TTFT 最大 9.39 倍短縮・SOTA 比 20.38% 精度改善を達成した。(Source: [[@2025__arXiv__KVShare - An LLM Service System with Efficient and Effective Multi-Tenant KV Cache Reuse]]) [[Dan Pei]](博士課程学生・指導教員の輩出元)が Build-bench 論文([[@2026__nkcs.iops.ai__Can Language Models Go Beyond Coding - Assessing the Capability of Language Models to Build Real-World Systems]], nkcs.iops.ai 2026-05)に共著者として参加。[[Nankai University]]・[[Peking University]]・[[Microsoft]] とのクロス ISA(x86_64/aarch64)ビルド修復ベンチマーク開発に寄与した。(Source: [[@2026__nkcs.iops.ai__Can Language Models Go Beyond Coding - Assessing the Capability of Language Models to Build Real-World Systems]]) [[Dan Pei]] が OScope 論文([[@2026__ICSE-SEIP__When LLMs Listen to Experts - Accurate Failure Diagnosis in Operating Systems]], ICSE-SEIP '26)に共著者として参加。[[Nankai University]]([[Yongxin Zhao]] 筆頭・[[Shenglin Zhang]] 責任著者)・[[Alibaba Group]] との共同で、OS 障害診断向け LLM フレームワーク [[OScope]] の開発に寄与した。(Source: [[@2026__ICSE-SEIP__When LLMs Listen to Experts - Accurate Failure Diagnosis in Operating Systems]]) [[Dan Pei]] が [[PerfScout]] 論文([[@2026__ICSE-SEIP__PerfScout - An Adaptive Workload Generator in Software Performance Testing]], ICSE-SEIP '26)に共著者として参加。[[Nankai University]]([[Yongqian Sun]]・[[Shenglin Zhang]] 責任著者ほか)・[[BizSeer]]([[Xidao Wen]])・[[Huawei Cloud]](成都)との共同で、SPOT(極値理論)・ADF/KPSS(局所定常性判定)・PPO(強化学習)を統合した性能テスト向け適応的ワークロード生成フレームワーク [[PerfScout]] の開発に寄与した。Huawei Cloud の CodeArts PerfTest に 9 か月間本番デプロイされ、ブレークポイント特定精度 82% 超を達成。RCA・異常検知が中心だった NetManAIOps グループの研究射程が性能テスト自動化にも広がったことを示す。(Source: [[@2026__ICSE-SEIP__PerfScout - An Adaptive Workload Generator in Software Performance Testing]]) [[Dan Pei]]・[[Qingyi Guo]] が OpsMem 論文([[@2026__arXiv__OpsMem - Dual-Memory Reasoning with Cross-Memory Resonance for Failure Diagnosis]], arXiv 2026-07)に共著者として参加。[[Nankai University]]([[Yongqian Sun]] 筆頭・[[Shenglin Zhang]] 責任著者ほか)・[[Huawei Technologies]] との共同で、短期記憶と長期記憶を cross-memory resonance で結合する失敗診断向けデュアルメモリフレームワーク [[OpsMem]] の開発に寄与した。(Source: [[@2026__arXiv__OpsMem - Dual-Memory Reasoning with Cross-Memory Resonance for Failure Diagnosis]]) [[Yihe Wang]](Department of Computer Science and Technology)が LLM ハルシネーション引用の大規模監査研究([[@2026__arXiv__LLM hallucinations in the wild]]、arXiv 2026-05-08)に共著者として参加。[[Cornell University]]・[[University of California, Berkeley]] のグループと共同で arXiv・bioRxiv・SSRN・PubMed Central 4 コーパスを横断監査し、2025 年単年で 146,932 件のハルシネーション引用を推定した。従来の AIOps/システム系ソースとは異なる、科学計量学(science of science)分野からの新規ソースである。(Source: [[@2026__arXiv__LLM hallucinations in the wild]]) [[Dan Pei]] が HeaRank 論文([[@2026__arXiv__Don't Predict, Prioritize - Rethinking GPU Reliability Assessment]]、KDD '26 V.2)の Computer Science Department 所属著者として参加。CNIC/CAS の [[Changhua Pei]]・[[Gaogang Xie]] らおよび [[StepFun]] の [[Yibo Zhu]] との共同で、GPU 障害の時系列予測が本質的に困難であることを実証し、Learning-to-Rank によるホストリスクランキングモデル HeaRank を提案した。マイクロサービス RCA・時系列異常検知が中心だった NetManAIOps グループの研究射程に、GPU ハードウェア信頼性という新ドメインが加わった。(Source: [[@2026__arXiv__Don't Predict, Prioritize - Rethinking GPU Reliability Assessment]]) ## 関連 - ソース: [[@2026__arXiv__Don't Predict, Prioritize - Rethinking GPU Reliability Assessment]] / [[@2026__nkcs.iops.ai__Can Language Models Go Beyond Coding - Assessing the Capability of Language Models to Build Real-World Systems]] / [[@2026__arXiv__OpsMem - Dual-Memory Reasoning with Cross-Memory Resonance for Failure Diagnosis]] / [[@2026__ICSE-SEIP__When LLMs Listen to Experts - Accurate Failure Diagnosis in Operating Systems]] / [[@2025__NSDI__Minder - Faulty Machine Detection for Large-scale Distributed Model Training]] / [[@2025__NeurIPS2025__STRATUS - A Multi-agent System for Autonomous Reliability Engineering of Modern Clouds]] / [[@2025__CSUR__A Survey of AIOps in the Era of Large Language Models]] / [[@2025__ICLR__OpenRCA - Can Large Language Models Locate the Root Cause of Software Failures]] / [[@2025__SIGCOMM__Hawkeye - Diagnosing RDMA Network Performance Anomalies with PFC Provenance]] / [[@2025__SIGCOMM__SkeletonHunter - Diagnosing and Localizing Network Failures in Containerized Large Model Training]] / [[@2024__PVLDB__D-Bot - Database Diagnosis System using Large Language Models]] / [[@2025__SIGMOD__Rethinking The Compaction Policies in LSM-trees]] / [[@2026__arXiv__Position - The Inevitable End of One-Architecture-Fits-All-Domains in Time Series Forecasting]] / [[@2022__ACL__GLM - General Language Model Pretraining with Autoregressive Blank Infilling]] / [[@2025__arXiv__KVShare - An LLM Service System with Efficient and Effective Multi-Tenant KV Cache Reuse]] / [[@2026__ICSE-SEIP__PerfScout - An Adaptive Workload Generator in Software Performance Testing]] - ノブチューニングサーベイ([[@2023__TKDE__Automatic Database Knob Tuning - A Survey]], IEEE TKDE 2023)の全著者 [[Xinyang Zhao]]・[[Xuanhe Zhou]]・[[Guoliang Li]] の所属。ノブチューニングのパイプラインを4段階に分解し、16手法を体系的に比較した初の包括的サーベイ。(Source: [[@2023__TKDE__Automatic Database Knob Tuning - A Survey]]) - エンティティ: [[Yangtao Deng]] / [[ByteDance]] / [[University of Illinois Urbana-Champaign]] / [[IBM Research]] / [[Peking University]] / [[Alibaba Group]] / [[SkeletonHunter]] / [[Xuanhe Zhou]] / [[Guoliang Li]] / [[DB-GPT]] / [[Hengrui Wang]] / [[Huanchen Zhang]] / [[EcoTune]] / [[Qinwei Ma]] / [[Jingzhe Shi]] / [[Zaiwen Yang]] / [[Xinyang Zhao]] / [[Nankai University]] / [[BizSeer]] / [[Huawei Cloud]] / [[PerfScout]] / [[Yihe Wang]] / [[Cornell University]] / [[University of California, Berkeley]] - 概念: [[定常性モデル]]