# Huawei Cloud Huawei のクラウド部門(Huawei Cloud Computing Technology Co., Ltd.)。Computing and Networking Innovation Lab を擁し、本番マルチテナント LLM 訓練基盤 [[Platform-X]] を運営する。(Source: [[@2025__ESEC-FSE__L4 - Diagnosing Large-scale LLM Training Failures via Automated Log Analysis]]) - [[GRLIA]] 論文([[@2021__ASE__Graph-based Incident Aggregation for Large-Scale Online Service Systems]], ASE 2021)では、Networking サービスの 2020 年 5〜11 月本番インシデントを提供し、インシデント管理システムへの GRLIA 展開を行った。論文は 2020 年 11 月の 26 障害で平均障害対応時間が過去 3 か月比 18.6〜24.8% 短縮したと報告する。(Source: [[@2021__ASE__Graph-based Incident Aggregation for Large-Scale Online Service Systems]]) - LLM 訓練障害のログ解析診断 [[L4]](FSE 2025)と、ネットワークフローからのブラックボックス性能診断 [[LLMPrism]](DSN 2025)の双方で産業側著者(Cong Feng・Yongqiang Yang・Zengyin Yang・[[Rui Ren]])を出す。いずれも [[The Chinese University of Hong Kong]] の [[Michael R. Lyu]] グループとの産学共同研究。 - [[L4]]・[[LLMPrism]] はともに本番の [[Platform-X]] にデプロイされ、L4 は 2024 年 6 月から障害管理システムに、LLMPrism は 2024 年 10 月から稼働する。 - [[FlowXpert]] 論文([[@2025__KDD__FlowXpert - Expertizing Troubleshooting Workflow Orchestration with Knowledge Base and Multi-Agent Coevolution]], KDD 2025)では Zhi Zhang・Ronghua Sun・Haihua Li(Huawei Cloud)と Wei Song・Xiaolong Chen・Jingbo Miao(Huawei Technologies Ltd.)が産業著者として参加。Alarmagnify システムに FlowXpert を統合し、DCN の 34,488 件インシデントで 10 週間展開した(承認率約 80%、生成時間 22.1 秒)。(Source: [[@2025__KDD__FlowXpert - Expertizing Troubleshooting Workflow Orchestration with Knowledge Base and Multi-Agent Coevolution]]) - INLG 2025 のサーベイ論文「Taming the Titans」では、[[Qingrong Xia]]・[[Xinyu Duan]]・[[Zhefeng Wang]]・[[Baoxing Huai]] が Huawei Cloud 所属として、LLM 推論サービングのモデル配置・スケジューリング・KV キャッシュ・[[Prefill-Decode分離]]・クラスタ配置・新興シナリオを整理した。(Source: [[@2025__INLG__Taming the Titans - A Survey of Efficient LLM Inference Serving]]) - [[PerfScout]] 論文([[@2026__ICSE-SEIP__PerfScout - An Adaptive Workload Generator in Software Performance Testing]], ICSE-SEIP '26)では、成都拠点の Huandong Zhuang・Bowen Deng・Ruiyuan Wan が産業側著者として参加し、SPOT・ADF/KPSS・PPO を統合した強化学習ベースの適応的ワークロード生成フレームワーク [[PerfScout]] をクラウドベース性能テストプラットフォーム CodeArts PerfTest に 9 か月間本番デプロイした。実運用データセット D1(56 件・データ/監視サービス 10 API)・D2(146 件・データ/監視/認証サービス 16 API)で評価し、ブレークポイント特定精度 82% 超・テスト効率改善 90% 近くを達成。L4・LLMPrism([[The Chinese University of Hong Kong]] との連携)が LLM 訓練基盤 Platform-X の障害診断・性能診断を担う一方、PerfScout は([[Nankai University]]・[[BizSeer]]・[[Tsinghua University]] との連携で)性能テストのワークロード生成自動化を担い、Huawei Cloud の AIOps 産学連携が複数の学術グループ・複数の運用課題にまたがることを示す。(Source: [[@2026__ICSE-SEIP__PerfScout - An Adaptive Workload Generator in Software Performance Testing]]) ## 関連 - ソース: [[@2025__ESEC-FSE__L4 - Diagnosing Large-scale LLM Training Failures via Automated Log Analysis]] / [[@2025__DSN__LLMPrism - Black-box Performance Diagnosis for Production LLM Training Platforms]] / [[@2025__KDD__FlowXpert - Expertizing Troubleshooting Workflow Orchestration with Knowledge Base and Multi-Agent Coevolution]] / [[@2025__INLG__Taming the Titans - A Survey of Efficient LLM Inference Serving]] / [[@2026__ICSE-SEIP__PerfScout - An Adaptive Workload Generator in Software Performance Testing]] - 所属研究者: [[Cong Feng]] / [[Yongqiang Yang]] / [[Zengyin Yang]] / [[Rui Ren]] / [[Qingrong Xia]] / [[Xinyu Duan]] / [[Zhefeng Wang]] / [[Baoxing Huai]] / Huandong Zhuang / Bowen Deng / Ruiyuan Wan - 共同研究先: [[The Chinese University of Hong Kong]] / [[Nankai University]] / [[BizSeer]] / [[Tsinghua University]] - 関連プロダクト: [[Platform-X]] / [[L4]] / [[LLMPrism]] / [[PerfScout]] - 概念: [[ログ解析]] / [[LLM学習モニタリング]] / [[根本原因分析]] / [[定常性モデル]]