# Entities Index
### 2026-07-21 Tales from the Lunar Module Guidance Computer ingest
- [[Don Eyles]](新規) — Apollo LGC フライトソフトウェアエンジニア([[MIT Instrumentation Laboratory]])。スロットル制御ルーチン担当。
- [[Allan Klumpp]](新規) — 月着陸誘導方程式の設計者。throttle castellation の解析者。
- [[Hal Laning]](新規) — Apollo Guidance Computer の Executive/Waitlist リアルタイムOS設計者。
- [[Apollo Guidance Computer]](新規) — Apollo宇宙船搭載のフライトコンピュータ(AGC/LGC)。
- [[MIT Instrumentation Laboratory]](新規) — Apollo PGNCS の設計・開発機関(後のCharles Stark Draper Laboratory)。
- [[Margaret Hamilton]](更新) — Executive設計の個人帰属粒度に関するcontradiction calloutを追加。
### 2026-07-18 OpsMem ingest-paper
- [[OpsMem]](新規) — STM/LTM を cross-memory resonance で結合する失敗診断向けデュアルメモリフレームワーク。Nankai/Tsinghua/Huawei。
- [[Rongchen Gao]](新規) — OpsMem 共著者(Nankai University)。GoS(ICML 2026)共著者でもある。
- [[Qingyi Guo]](新規) — OpsMem 共著者(Tsinghua University)。
- [[Yaoliang Wu]](新規) — OpsMem 共著者(Huawei Technologies)。
- [[Yongqian Sun]](更新) — OpsMem 第一著者として追記。
- [[Yu Luo]](更新) — OpsMem 共著者として追記。GoS 筆頭著者との関連も明記。
- [[Wenwei Gu]](更新) — OpsMem 共著者(Nankai)として追記、9本目の Nankai 側継続共著。
- [[Shenglin Zhang]](更新) — OpsMem 責任著者として追記。
- [[Dan Pei]](更新) — OpsMem 共著者として追記。
- [[Qiuai Fu]](更新) — OpsMem 共著者(Huawei Technologies)として追記。
- [[Nankai University]](更新) — OpsMem 開発元として追記。
- [[Tsinghua University]](更新) — OpsMem 共同開発機関として追記。
- [[Huawei Technologies]](更新) — OpsMem 産業側共同開発機関・データセット提供元として追記。
### 2026-07-18 MLCommons Chakra ingest-paper
- [[MLCommons Chakra]](新規) — Chakra Execution Trace(ET)とその収集・分析・リプレイ・シミュレーションエコシステム。MLCommonsのワーキンググループとして標準化。
- [[MLCommons]](新規) — AI/ML性能ベンチマーキング標準化団体。MLPerfで知られる。
- [[Georgia Institute of Technology]](新規) — MetaとChakraを共同立ち上げ。ASTRA-sim開発機関。
- [[Tushar Krishna]](新規) — Georgia Tech准教授。Chakra論文の責任著者の一人。
- [[Srinivas Sridharan]](新規) — NVIDIA所属。Chakra論文の責任著者の一人。
- [[ASTRA-sim]](新規) — Chakra ETをネイティブサポートするオープンソース分散AIシステムシミュレータ。
- [[NVIDIA]](更新) — Chakraワーキンググループへの主要な貢献企業として追記。
- [[AMD]](更新) — Chakra ETのJSON形式追加とASTRA-sim両対応化を追記。
- [[vLLM]](更新) — Chakraのinferenceトレース統合(MoEルーティング・KVオフロード・PD分離KV転送分析)を追記。
### 2026-07-15 Scalable and Energy-Efficient AI ingest-paper
- [[Muhammad Ali Shafique]](新規) — Kansas State University所属。筆頭著者。
- [[Imran Latif]](新規) — Johnson Controls所属。GPUノード電力実測の先行研究にも関与。
- [[Hayat Ullah]](新規) — Florida Atlantic University所属。
- [[Alex C. Newkirk]](新規) — Lawrence Berkeley National Laboratory所属。GPUノード電力実測の先行研究にも関与。
- [[Arslan Munir]](新規) — Florida Atlantic University所属。corresponding author。
- [[Kansas State University]](新規) — Manhattan, KS, USA所在の州立大学。
- [[Johnson Controls]](新規) — Milwaukee, WI, USA所在の企業。
- [[Florida Atlantic University]](新規) — Boca Raton, FL, USA所在の州立大学。
- [[Lawrence Berkeley National Laboratory]](新規) — Berkeley, CA, USA所在の米国エネルギー省国立研究所。
### 2026-07-15 Speculations Concerning the First Ultraintelligent Machine ingest-paper
- [[I. J. Good]](更新) — 一次論文本体([[@1965__AdvComput__Speculations Concerning the First Ultraintelligent Machine]])を追加し、「未 ingest」注記を解消。役割説明を更新。
### 2026-07-14 言語モデルの内部機序:解析と解釈 ingest-slides
- [[Benjamin Heinzerling]](新規) — 理化学研究所・東北大学所属。言語モデルの内部表現の機構的解釈性研究者。
- [[横井祥]](新規) — 国立国語研究所・東北大学・理化学研究所所属。
- [[小林悟郎]](新規) — 東北大学・理化学研究所所属。
- [[理化学研究所]](新規) — 日本の公的研究機関。登壇者3名の所属機関の一つ。
- [[東北大学]](新規) — 日本の国立大学。登壇者3名の所属機関の一つ。
- [[国立国語研究所]](新規) — 日本語を対象とする公的研究機関。横井祥の所属機関の一つ。
- [[Anthropic]](更新) — SAEによるゴールデンゲートブリッジ特徴の事例を追記。
### 2026-07-10 Failure Trends in a Large Disk Drive Population ingest
- [[Eduardo Pinheiro]](新規) — [[Google]] Inc. エンジニア。大規模 HDD 障害傾向分析(FAST 2007)の筆頭著者。(person / storage / reliability)
- [[Wolf-Dietrich Weber]](新規) — [[Google]] Inc. エンジニア。大規模 HDD 障害傾向分析(FAST 2007)の共著者。(person / storage / reliability)
- [[Luiz André Barroso]](更新) — FAST 2007 HDD 信頼性研究の共著者として追記。
### 2026-07-08 ICPE Benchmarking ingest
- [[David Georg Reichelt]](新規) — Lancaster University Leipzig / URZ Leipzig 所属。[[MooBench]] 主要開発者。Java トレーシングエージェントオーバーヘッド研究。
- [[Wilhelm Hasselbring]](新規) — Christian-Albrechts-Universität zu Kiel 教授。[[Kieker]] 主要開発者。
- [[MooBench]](新規) — Java オブザーバビリティフレームワークの計装オーバーヘッド計測マイクロベンチマーク。
- [[Kieker]](新規) — IRL ベース拡張可能な Java 監視フレームワーク。7 フレームワーク中最低オーバーヘッド。
### 2026-07-06 ARGUS ingest
- [[Jiasheng Zhou]](新規) — [[Tencent]] 所属、10,000 GPU 超クラスター向け常時稼働トレーシング・診断システム ARGUS の筆頭著者。
- [[Longbin Zeng]]・[[Clavis Chen]]・[[Ruiming Lu]]・[[Qinwei Yang]]・[[Leyi Ye]]・[[Ray Ying]]・[[Key Zhang]](新規) — Tencent 所属の ARGUS 共著者群。
- [[Tencent]](更新) — ARGUS 開発チームの所属組織として追記。
### 2026-07-06 KRCA ingest
- [[Jiamin Jiang]](新規) — [[Nankai University]] 所属、[[Kuaishou Technology]] インターン中に KRCA 開発した第一著者。
- [[Yongqian Sun]]・[[Dan Pei]]・[[Kuaishou Technology]]・[[Nankai University]](更新) — KRCA 共著者および所属機関として追記。
### 2026-07-07 VAST AI Operating System ingest
- [[VAST Data]](更新) — VAST AI OS の主要コンポーネント(DASE・DataStore・DataBase・DataSpace・DataEngine・InsightEngine)を追記。ソース追加。
### 2026-07-06 A Multi-Dataset Benchmark for Evaluating LLM Agents in Microservice Failure Diagnosis ingest
- [[Yuanhong Cai]](新規) — CNIC/CAS 所属研究者・第一著者。
- [[Changhua Pei]]・[[Dan Pei]]・[[Shenglin Zhang]]・[[Yongqian Sun]]・[[Xiaohui Nie]]・[[Kanglin Yin]]・[[Xidao Wen]](更新) — 本論文の共著者として追記。
### 2026-07-06 A Checkpoint/Restore Mechanism with Interoperability Among Distinctive WebAssembly Interpreters (APSys 2024 Poster) ingest
- [[Wasm3]](新規) — fast interpreter の代表例の一つ。WAMR と並んで APSys 2024 で異種 interpreter 間 C/R の対象ランタイムとして扱われる。
- [[Daigo Fujii]]・[[Katsuya Matsubara]]・[[Yuki Nakata]]・[[Future University Hakodate]]・[[SAKURA internet Inc.]]・[[WasmEdge]]・[[WAMR]](更新) — 著者、所属機関、対象ランタイムとして追記。
### 2026-07-05 Self-Hosted WebAssembly Runtime for Runtime-Neutral Checkpoint/Restore in Edge–Cloud Continuum (Mid4CC ’25) ingest
- [[Chiwawa]]・[[Wizard]]・[[CRIU]](新規) — 自己ホスト型 Wasm ランタイムのプロトタイプ、比較対象の自己ホスト可能ランタイム、チェックポイントサイズ比較基準の Linux プロセス C/R ツール。
- [[Yuki Nakata]]・[[Katsuya Matsubara]]・[[Future University Hakodate]]・[[SAKURA internet Inc.]]・[[WasmEdge]]・[[WAMR]](更新) — 著者、所属機関、評価対象ホストランタイムとして追記。
### 2026-07-05 Seamless Self-Healing in WebAssembly Container Orchestration with Runtime-Neutral Checkpointing (CANDARW 2025) ingest
- [[Yuzuki Saito]](新規) — CANDARW 2025 論文の共著者。
- [[Katsuya Matsubara]]・[[Daigo Fujii]]・[[Yuki Nakata]]・[[Future University Hakodate]]・[[SAKURA internet Inc.]](更新) — 著者、所属機関として追記。
- [[WasmEdge]]・[[WAMR]](更新) — 動的ランタイム切り替えの対象ランタイムとして追記。
### 2026-07-05 Container Transplantation (APSys ’23) ingest
- [[Shintaro Suzuki]](新規) — Container Transplantation 論文の共著者。
- [[gVisor]]・[[Kata Containers]]・[[FreeBSD]]・[[Linux]]・[[Linuxulator]](新規) — 比較対象または実行基盤となるシステム。
- [[Yuki Nakata]]・[[Katsuya Matsubara]]・[[SAKURA internet Inc.]]・[[Future University Hakodate]]・[[Docker]](更新) — 著者、所属機関、関連製品として追記。
### 2026-07-05 Stateful VM Migration Among Heterogeneous WebAssembly Runtimes (EdgeSys ’24) ingest
- [[Daigo Fujii]]・[[WasmEdge]]・[[WAMR]](新規) — EdgeSys ’24 論文の筆頭著者および対象ランタイム。
- [[Yuki Nakata]]・[[Katsuya Matsubara]]・[[Future University Hakodate]]・[[SAKURA internet Inc.]](更新) — 著者、所属機関として追記。
### 2026-07-05 Subaco (UCC 2021) ingest
- [[Yuki Nakata]]・[[Katsuya Matsubara]]・[[Ryosuke Matsumoto (SAKURA internet)|Ryosuke Matsumoto]](新規) — Subaco 論文の著者。
- [[Future University Hakodate]]・[[SAKURA internet Inc.]](新規) — 著者の所属機関。
### 2026-07-04 OSDI'25 Extending Applications Safely and Efficiently ingest
- [[Yanpeng Hu]]・[[Xiaozheng Lai]]・[[Dan Williams]]・[[Andi Quinn]]・[[Redis]]・[[FUSE]]・[[OpenSSL]](新規) — bpftime/EIM 論文の共著者またはユースケース対象。
- [[Yusheng Zheng]]・[[Tong Yu]]・[[Yiwei Yang]]・[[bpftime]]・[[eunomia-bpf]]・[[DeepFlow]]・[[Nginx]](更新) — OSDI'25 ソースを関連ソースとして追記。
### 2026-07-04 The GPU Observability Gap ingest
- [[Yusheng Zheng]]・[[Tong Yu]]・[[Yiwei Yang]]・[[eunomia-bpf]]・[[bpftime]](更新) — eGPU ブログ記事を関連ソースとして追記。
### 2026-07-04 CUDA Events - eBPF-based CUDA API Tracing ingest
- [[yunwei37]](新規) — [[eunomia-bpf]] コミュニティの中心的人物。eBPF チュートリアル群(CUDA Events 等)と bpftime/eGPU 関連プロジェクトの主要コントリビューター。(person / ebpf)
- [[eunomia-bpf]]・[[bpftime]]・[[libbpf]]・[[NVIDIA]](更新) — CUDA API トレースチュートリアルを関連ソース・プロジェクトとして追記。
- [[null2]](新規) — 大阪・関西万博2025における [[計算機自然]] の公共規模インスタレーション。鏡面建築、身体のデジタル化、生成AI、ロボティックな変形、null/空のモチーフを統合する。(product / media-art)
- [[xDiversity]](新規) — 身体能力の違いへ応答する計算的に多様な社会実装プロジェクト。[[アクセシビリティ]] を計算機自然の存在論的試金石にする実践例。(organization / accessibility)
- [[Digital Nature Group]](新規) — [[落合陽一]]と結びつく研究・制作グループ。音響操作、空中グラフィクス、[[xDiversity]]、[[null2]] などを通じて [[計算機自然]] を実践的に展開する。(organization / media-art)
- [[落合陽一]](更新) — [[デジタル発酵]]・[[デジタル蒸留]]・[[Homo Convivium]] を生成AI以後の後期語彙として追記。
### 2026-07-04 計算機自然からマタギドライヴへ ingest
- [[落合陽一]](新規) — メディアアーティスト・研究者。[[計算機自然]]の提唱者であり、[[マタギドライヴ]]、[[主体なき美の美学]]、[[ヌルのテトラレンマ]]を通じて計算・物質・身体・自然概念を横断的に論じる。(person / media-art / philosophy-of-technology)
### 2026-07-02 PLaMo 2 Technical Report ingest
- [[Preferred Networks]](新規) — 日本の AI 企業。PLaMo 2 Technical Report の著者組織であり、[[PLaMo 2]] を開発。(organization / llm)
- [[PLaMo 2]](新規) — Preferred Networks の日本語重視 LLM 系列。Samba ベース構成、構造化枝刈り、事後学習、vLLM 推論最適化を特徴とする。(product / llm)
### 2026-07-02 XProf (MLSys 2026) ingest
- [[Robert Hundt]](新規) — [[Google]] Cloud シニアエンジニア。XProf の主要設計者・MLSys 2026 筆頭著者。(person / ml-systems)
- [[OpenXLA]](新規) — Google 主導のオープンソース ML コンパイラ・ランタイムエコシステム。XLA コンパイラ・PJRT・XProf を含む。(organization / ml-systems)
- [[Google]](更新) — XProf の開発元として [[MLプロファイリング]] セクションを追加。
### 2026-07-02 The Case for Learned Index Structures ingest
- [[Alex Beutel]](新規) — [[Google]] 所属の研究者。Learned Index Structures 論文の共著者。(person / database)
- [[Ed H. Chi]](新規) — [[Google]] 所属の研究者。Learned Index Structures 論文の共著者。(person / database)
- [[Neoklis Polyzotis]](新規) — [[Google]] 所属の研究者。Learned Index Structures 論文の共著者。(person / database)
- [[Tim Kraska]](更新) — [[Learned Index]] の提案ソースを追記。Kraska の貢献は Google 所属中の研究と注記される。(person / database / machine-learning)
- [[Jeffrey Dean]](更新) — Learned Index Structures への共著を追記。(person / google / database)
- [[Google]](更新) / [[MIT]](更新) — 著者所属と研究背景を追記。(organization / database)
#### Modernizing Incident Response with LLMs, RAG, and the MCP (SREcon25 EMEA, 2025) (2026-07-01)
- [[Theofilos Papapanagiotou]](新規) — [[Amazon]] の SRE インフラ担当エンジニア。MCP + RAG による障害対応刷新を SREcon25 EMEA で発表。(person / sre / amazon)
- [[Amazon]](更新) — MCP + RAG による障害対応の産業実装節を追加。(organization / distributed)
- [[Model Context Protocol]](更新) — Amazon の人間/エージェント共通ハンドルと IAM ロール分離認証の産業実装例を追加。(product / protocol / agent)
### 2026-07-01 The Un-Incident (SREcon25 EMEA) ingest
- [[Andreas Deuschl]](新規) — [[Dynatrace]] プロダクトリード(Delivery / Reliability / Security)。SRE・Ops・セキュリティ 25 年超のプラクティショナー。Un-Incident 概念と Gray Zone Playbook の提唱者。(person / sre / incident-management)
- [[Dynatrace]](新規) — オーストリア発祥のオブザーバビリティ・AIOps・セキュリティプラットフォーム企業。Davis CoPilot(AI 支援インシデント対応)を提供。(organization / observability / sre)
### 2026-07-01 Incident Groundhog Day (SREcon24 EMEA) ingest
- [[Hamed Silatani]](新規) — [[Uptime Labs]] 共同創業者 & CEO。ステージドワールド手法を用いたインシデントシミュレーション実験の設計者。(person / sre / incident-simulation)
- [[Uptime Labs]](新規) — インシデント学習・シミュレーションプラットフォーム企業。AWS+Kubernetes 上にリアルな eコマース環境を構築し、AI ボットチームと Slack ブリッジを組み合わせたステージドワールドを提供する。(organization / sre / incident-simulation)
### 2026-07-01 Incident Management Metrics that Matter (SREcon25 Americas) ingest
- [[Jamie Luck]](新規) — [[Datadog]] シニア SRE。インシデント管理チームの暫定マネージャーを経験。代名詞 they/them。(person / sre)
- [[Laura de Vesine]](更新) — 役職を「スタッフエンジニア」から「シニアスタッフエンジニア」に更新、SREcon25 Americas 発表を追加。(person / sre)
- [[Datadog]](更新) — SREcon25 Americas 発表を出典リストに追加。(organization / sre / aiops)
### 2026-07-01 Embracing the Multi-Party Dilemma (SREcon23 EMEA) ingest
- [[Alex Elman]](新規) — [[Indeed]]。Learning from Incidents 実践・Multi-Party Dilemma 研究の共同提唱者。(person / sre / incident-response)
- [[SentinelOne]](新規) — [[Sarah Butt]] の登壇時点(2023年)の所属先。(organization / security)
- [[Sarah Butt]](更新) — Salesforce SRE(2021年時点)から SentinelOne SRE(2023年時点)への異動を反映。(person / sre)
- [[Indeed]](更新) — Learning from Incidents 実践と Multi-Party Dilemma 発見の経緯を追記。(organization / sre)
- [[Laura Maguire]](更新) — 協調コスト(coordination cost)論点の接続を追記。(person / resilience-engineering / sre)
- [[David D. Woods]](更新) — Multi-Party Dilemma という語の共同提唱者として追記。(person / resilience-engineering)
- [[John Allspaw]](更新) — 同上、David D. Woods との共同提唱を追記。(person / sre)
- [[Richard I. Cook]](更新) — 謝辞記載として出典を追加。(person / human-factors)
### 2026-07-01 An Organizational Response to Incidents (SREcon23 Americas) ingest
- [[Laura Maguire]](更新) — SREcon23 Americas 単独登壇。フォロワーシップ(Followship)を提唱。(person / resilience-engineering / sre)
- [[Jeli]](更新) — People View・Narrative Marker 等の実プロダクト画面が事例として使われる。(organization / sre)
### 2026-07-01 Handover Communications in Software Operations (SREcon23 Americas) ingest
- [[Chad Todd]](新規) — CrowdStrike のエンジニア。Lund University 大学院で人的要因・安全科学を専攻。SREcon23 Americas 登壇者。(person / sre / human-factors)
- [[CrowdStrike]](新規) — サイバーセキュリティ企業。Network Operations Center・Customer Support Center が引き継ぎ伝達研究の調査対象。(organization / company)
- [[Lund University]](新規) — スウェーデンの大学。人的要因・安全科学の大学院課程。(organization / university)
- [[David D. Woods]](新規) — Ohio State University。レジリエンスエンジニアリング・認知システム工学の中核的研究者。(person / resilience-engineering / human-factors)
- [[Emily Patterson]](新規) — 引き継ぎコミュニケーション研究史における中核的研究者の一人。(person / human-factors)
- [[Gary Klein]](新規) — 認知科学者。Joint Activity・Common Ground 概念の原典に位置づけられる。(person / human-factors / resilience-engineering)
### 2026-07-01 The Math behind the Incident Aftermath (SREcon22 APAC) ingest
- [[Ashish Patel]](新規) — PayPal Site Reliability Platform Engineering ソフトウェアエンジニア。SREcon22 APAC 登壇者。(person / sre)
- [[Sriram Srinivasan]](新規) — PayPal Technical Architect。SRE 領域で20年の経験。SREcon22 APAC 登壇者。(person / sre)
- [[PayPal]](新規) — 決済プラットフォーム企業。インシデント影響測定ツールを内製し SREcon22 APAC で発表。(organization / fintech)
### 2026-07-01 Evolution of Incident Management at Slack (SREcon21) ingest
- [[Brent Chapman]](新規) — Slack Staff Engineer(Reliability Pillar)。元 Google SRE、iMAG(Incident Management At Google)の設計者。SREcon21登壇者。(person / sre / incident-management)
- [[Slack Technologies]](更新) — 2018年9月 reliability crisis と Incident Management プログラム構築の経緯を追加。(organization / incident-management)
- [[PagerDuty]](更新) — Slack の IC/Responder 訓練クラスの土台となった公開クラスの言及を追加。(organization / incident-management)
### 2026-07-01 You Can't Stop Fires with an Ambulance (SREcon18 Asia) ingest
- [[Piers Chamberlain]](新規) — Xero の Site Reliability Engineering 責任者。2011年入社。SREcon18 Asia 登壇者。(person / sre)
- [[Xero]](新規) — クラウド会計 SaaS 企業。2006年創業、2016年クラウド移行を機に SRE チーム発足。(organization / saas)
- [[Klaxon]](新規) — Xero の症状ベースアラートシステム。顧客ページヒット率を SQS+Lambda+SumoLogic+DataDog で検知。(product / alerting)
- [[Multivac]](新規) — Xero のインシデント管理 chatbot。war room を代替し post-mortem を自動生成。(product / chatops / incident-management)
- [[Report Card]](新規) — Xero の運用衛生スコアリングツール。SLO・アラート疲労等を採点し経営層との対話を促す。(product / slo)
### 2026-07-01 Fixing On-Call When Nobody Thinks It's (Too) Broken (SREcon19 Americas) ingest
- [[Tony Lykke]](新規) — Hudson River Trading Trade Systems SRE。以前はAtlassian (Sydney) でService Ops/Jira SRE。SREcon19 Americas登壇者。(person / sre)
- [[Hudson River Trading]](新規) — 自動化された取引会社(NYC)。米国株式日次出来高の5%超、~300人・10オフィス。Tony Lykkeが在籍するTrade Systems SREチームの母体。(organization / sre)
### 2026-07-01 Unified Theory of SRE (SREcon22 EMEA) ingest
- [[Emil Stolarsky]](新規) — SRE エンジニア。Wave Mobile Money 在籍(2022)。SREcon22 EMEA 登壇者。Seeking SRE 寄稿・97 Things Every SRE Should Know 共著。(person / sre)
- [[Wave Mobile Money]](参照) — アフリカのモバイル送金サービス。Emil Stolarsky が 2022 年時点で在籍していた組織。
### 2026-07-01 nrrd 911 ic me (SREcon16) ingest
- [[Alice Goldfuss]](新規) — New Relic SRE。@alicegoldfuss。ICS 社内展開推進者。SREcon16 Americas 登壇者。(person / sre)
### 2026-07-01 Software Engineering (Boehm 1976) ingest
- [[Barry W. Boehm]](新規) — TRW Systems and Energy Group ソフトウェア研究技術部長(1976 年時点)。ソフトウェアエンジニアリング定義・COCOMO・SREP の開発者。(person / software-engineering / classic)
- [[TRW Systems and Energy Group]](新規) — Redondo Beach, CA の防衛・航空宇宙企業。SREP/RSL/REVS・TRW ソフトウェア開発方法論の開発元。(organization / software-engineering)
### 2026-07-01 Notes from Production Engineering (SREcon15) ingest
- [[Pedro Canahuati]](新規) — Facebook Production Engineering ディレクター。SREcon15 登壇者。2009〜2015 年の組織変革を主導。(person / sre / facebook)
- [[Jay Parikh]](新規) — Facebook エンジニアリング責任者(Head of Engineering)。Production Engineering 改革の後ろ盾。週次 SEV レビューに同席し即座にリソース決定を行う。(person / sre / facebook)
- [[Facebook]](更新) — Production Engineering セクション追加(SRO/AppOps/PE 組織史・FBAR・Cobalt・ODS・週次 SEV レビュー)。
### 2026-06-30 Software Analytics for Incident Management ASE 2013 ingest
- [[Rui Ding]](新規) — Microsoft Research Asia(北京)の研究者。SAS 開発の共著者。(person / aiops)
- [[Qiang Fu]](新規) — Microsoft Research Asia(北京)の研究者。CAR マイニング・治癒行動推薦の共著者。(person / aiops)
- [[Tao Xie]](新規) — UIUC 教授。ソフトウェアアナリティクス研究者。SAS の学術側共著者。(person / software-engineering)
- [[Jian-Guang Lou]](更新) — first_mentioned を ASE 2013 SAS 論文に更新、共著者リストと関連概念を拡充。(person / aiops)
- [[Qingwei Lin]](更新) — ASE 2013 SAS 論文を sources に追加、本文に SAS 開発参加の記述を追加。(person / aiops)
- [[Dongmei Zhang]](更新) — ASE 2013 SAS 論文を sources に追加、本文に SAS 開発参加の記述を追加。(person / aiops)
### 2026-06-30 Xpert ICSE 2024 ingest
- [[Zhihao Yang]](新規) — [[Peking University]] 所属。Microsoft Research Asia インターン。Xpert (ICSE 2024) 共著者。(person / aiops)
- [[Yuxuan Jiang]](更新) — Xpert (ICSE 2024) での Microsoft インターン中の研究を追記。(person / aiops)
### 2026-06-30 Fail through the Cracks EuroSys 2023 ingest
- [[Lilia Tang]](新規) — UIUC 博士課程学生。CSI 障害研究の共同筆頭著者。(person / reliability)
- [[Chaitanya Bhandari]](新規) — UIUC 博士課程学生。CSI 障害研究の共同筆頭著者・クロスシステムテストのケーススタディを担当。(person / reliability)
- [[Indranil Gupta]](新規) — UIUC 教授。分散システム信頼性研究者。CSI 障害研究の共同 PI。(person / distributed)
- [[Tianyin Xu]](更新) — EuroSys '23 CSI 障害研究の最終著者として追記。(person / reliability)
- [[Purdue University]](更新) — Yongle Zhang の所属機関としてCSI 障害研究への参加を追記。(organization)
### 2026-06-30 Metastable Failures HotOS 2021 ingest
- [[Nathan Bronson]](新規) — Rockset(旧 Facebook)エンジニア。メタ安定障害体系化・リンク不均衡障害の解明に関与。(person / distributed-systems)
- [[Abutalib Aghayev]](新規) — Penn State 研究者。メタ安定障害 HotOS 2021 共著者。(person / distributed-systems)
- [[Aleksey Charapko]](新規) — University of New Hampshire 研究者。メタ安定障害 HotOS 2021 共著者。(person / distributed-systems)
- [[Timothy Zhu]](新規) — Penn State 研究者。メタ安定障害 HotOS 2021 共著者。(person / distributed-systems)
- [[Rockset]](新規) — Nathan Bronson らが設立したリアルタイム分析データベース企業。(organization)
- [[The Pennsylvania State University]](新規) — 米国ペンシルバニア州立大学(Penn State)。Aghayev・Zhu の所属機関。(organization)
- [[University of New Hampshire]](新規) — 米国ニューハンプシャー大学(UNH)。Charapko の所属機関。(organization)
### 2026-06-30 Gray Failure HotOS 2017 ingest
- [[Jacob R. Lorch]](新規) — Microsoft Research 研究者。Gray Failure HotOS 2017 共著者。(person / reliability)
- [[Murali Chintalapati]](新規) — Microsoft Azure エンジニア。Gray Failure HotOS 2017 共著者。(person / reliability)
- [[Randolph Yao]](新規) — Microsoft Azure エンジニア。Gray Failure HotOS 2017 共著者。(person / reliability)
- [[Peng Huang]](更新) — MSR→Johns Hopkins→U-Michigan の経歴と HotOS 2017 主著者としての役割を追記。
- [[Chuanxiong Guo]](更新) — Gray Failure HotOS 2017 共著者として role と貢献を追記。
- [[Lidong Zhou]](更新) — Gray Failure HotOS 2017 共著者リンクと追加出典を追記。
- [[Yingnong Dang]](更新) — Gray Failure HotOS 2017 共著者として role と出典を追記。
- [[Johns Hopkins University]](更新) — stub から実体に格上げ。Peng Huang の兼任所属として記述。
### 2026-06-30 mTCP NSDI 2014 ingest
- [[EunYoung Jeong]](新規) — KAIST 研究者。mTCP 筆頭著者。(person / networking)
- [[Dongsu Han]](新規) — KAIST 研究者。mTCP 共著者。(person / networking)
- [[KyoungSoo Park]](更新) — mTCP ソース参照とリンクを追記。
- [[KAIST]](更新) — EunYoung Jeong・Dongsu Han・mTCP ソース参照を追記。
### 2026-06-30 ISPASS 2015 — VM vs Linux Containers
- [[Wes Felter]](新規) — IBM Research Austin の研究者。ISPASS 2015「VM vs Containers」論文筆頭著者。
- [[Docker]](更新) — ISPASS 2015 の実測性能データ(NAT/AUFS のコスト、ランダム I/O 性能)を追記。
- [[IBM Research]](更新) — Austin TX の仮想化性能研究(ISPASS 2015)セクションを追記。
### 2026-06-30 NSDI 2013 Scaling Memcache at Facebook
- [[Rajesh Nishtala]](新規) — Facebook エンジニア。NSDI '13「Scaling Memcache at Facebook」筆頭著者。
- [[Facebook]](更新) — memcache(NSDI 2013)セクションと関連リンクを追記。
### 2026-06-30 LISA 2013 Merlin 論文
- [[Marc Merlin]](新規) — Google エンジニア。ProdNG ライブアップグレードを主導。
- [[Richard Gooch]](新規) — Google エンジニア。ProdNG 設計者。ELF バイナリパッチアイデアの考案者。
- [[Google]](更新) — インフラストラクチャ管理セクションと LISA '13 参照を追記。
### 2026-06-30 ERA 論文(Nature 2026)
- [[Michael P. Brenner]](新規) — [[Google Research]] / [[Harvard University]] 教授。ERA 論文の共同責任著者。
- [[DeepMind]](更新) — ERA 論文で Google DeepMind モントリオール著者群(Aygün, Comanici, Smalling, Mourad ほか)を追記。
- [[Google Research]](更新) — ERA 論文で [[Michael P. Brenner]] と ERA 関連ソース・概念を追記。
Navigation: [[index]] | [[concepts/_index]] | [[sources/_index]]
ソースに登場する実体(人物・組織・システム・データセット・プロジェクト)ページの一覧。
---
### 2026-06-30 Towards end-to-end automation of AI research (Nature 2026) ingest
- [[Chris Lu]](新規) — Sakana AI / FLAIR, University of Oxford。The AI Scientist 共同筆頭著者。テンプレートベース版開発をリード。(person / ai-research)
- [[Cong Lu]](新規) — Sakana AI / UBC / Vector Institute。The AI Scientist 共同筆頭著者。アイデア生成・自動査読者・IRB承認を主導。(person / ai-research)
- [[Robert Tjarko Lange]](新規) — Sakana AI。The AI Scientist 共同筆頭著者。自動査読者のコア実装・ワークショップ対応を主導。(person / ai-research)
- [[Yutaro Yamada]](新規) — Sakana AI(責任著者)。The AI Scientist 共同筆頭著者。ツリー探索コアとテンプレート自由版を実装。(person / ai-research)
- [[Shengran Hu]](新規) — Sakana AI / UBC / Vector Institute。VLM統合・自動査読者ベンチマークを担当。(person / ai-research)
- [[Jakob Foerster]](新規) — FLAIR, University of Oxford。The AI Scientist 共著。(person / ai-research)
- [[David Ha]](新規) — Sakana AI 創業者・責任著者。元 Google Brain。プロジェクト全体を統括。(person / ai-research)
- [[Jeff Clune]](新規) — University of British Columbia / Vector Institute(責任著者)。オープンエンドアルゴリズム研究の第一人者。IRB承認・AI論文評価を統括。(person / ai-research)
- [[Sakana AI]](更新) — The AI Scientist プロダクト追記。
### 2026-06-30 電動マイクロモビリティのシェアサービス「LUUP」におけるEnabling SLOの実践 (SRE NEXT 2023) ingest
- [[Wataru Tsuda]](新規) — Luup SRE エンジニア(gr1m0h)。SRE NEXT 2023 Chair。Enabling SLO と CMC 概念を主導。(person / sre / slo / iot)
- [[Luup]](新規) — 電動マイクロモビリティのシェアサービス「LUUP」を運営するスタートアップ。IoT × SRE の実践事例。(organization / iot / sre)
### 2026-06-30 Who owns the Service Level? (SRE NEXT 2022) ingest
- [[近藤武士]](新規) — Recruit Engineering Manager, SRE。スタディサプリ小中高プロダクト開発部 SRE チームを率いる。@chaspy。(person / sre / slo)
- [[Recruit]](新規) — リクルート(Recruit Co., Ltd.)。旧 Quipper Japan を統合しスタディサプリを運営。(organization / edtech)
- [[スタディサプリ]](新規) — Recruit が運営する日本向け教育 SaaS。2022 年時点で開発者 114 名・SRE 7 名体制。(product / edtech / sre)
### 2026-06-30 プロダクトオーナーとしてSLOに向き合う 〜Mackerelチームの事例〜 (SRE NEXT 2023) ingest
- [[渡辺 起]](新規) — 株式会社はてな Mackerel プロデューサー(2022 年まで PO)。SRE NEXT 2023 で Mackerel チームの SLO 実践を PO 視点で発表。(person / sre / slo)
- [[Mackerel]](更新) — SLO 導入事例を追記。チーム 10 人前後(SRE 1〜3 名)、DORA フロー/後期段階、Error Budget Policy は最低限アクションから開始。
### 2026-07-01 From 4 Hours to 8 Minutes with AI Agents that Transform SRE Incident Response (SREcon25 EMEA) ingest
- [[Peter Jausovec]](新規) — [[Solo.io]] エンジニア、[[kagent]] メンテナ。Kubernetes・Istio・AI インフラを専門。USENIX SREcon25 EMEA で AIRE フレームワークを発表。(person / sre / cloud-native / agent)
- [[Solo.io]](新規) — クラウドネイティブネットワーキング・AI インフラ企業。Kubernetes・Istio・サービスメッシュ専門。kagent を主導。(organization / cloud-native)
- [[kagent]](新規) — Kubernetes ネイティブ AI エージェント・MCP サーバ構築フレームワーク。CNCF サンドボックスプロジェクト。[[Peter Jausovec]] がメンテナ。(product / cloud-native / agent / mcp)
### 2026-06-30 Is the S in SRE for "Security"? (SREcon25 Americas) ingest
- [[John Benninghoff]](新規) — Security Differently 創業者。サイバーセキュリティと SRE の統合を安全科学(Safety-II)で論じる研究者・実践者。Trinity College Dublin で安全科学修士。(person / security / sre)
- [[Security Differently]](新規) — Benninghoff 設立のセキュリティコンサルティング。セキュリティを工学に統合するアドバイスを専門とする。(organization / security)
### 2026-06-30 Beyond Sequential (SREcon25 Americas) ingest
- [[Jash Mistry]](新規) — eBay Member of Technical Staff(SRE)。Georgia Tech 修士。非同期ランキングサービスの SLO 実装担当。(person / sre)
- [[Gabriela Medvetska]](新規) — eBay Site Reliability Engineer。UC Santa Cruz 学士。マルチウィンドウアラート設計担当。(person / sre)
- [[eBay]](更新) — 非同期ランキングサービスへの SLO 適用(ML 推薦導入で旧システム比 2.1% 収益増)を追記。
### 2026-06-30 9 Things You Should Do When Starting to Use SLOs (SREcon23 EMEA) ingest
- [[Sal Furino]](新規) — Customer Reliability Engineer。SLO 導入・SRE 文化浸透を専門。SLODLC フレームワークを紹介。(person / sre)
### 2026-06-30 SLX: An Extended SLO Framework (SREcon21) ingest
- [[Qian Ding]](新規) — Ant Group スタッフエンジニア・SRE テックリード。60+ K8s クラスタの Infra SRE と SLX フレームワーク設計者。(person / sre)
- [[Xuan Zhang (Ant Group)]](新規) — Ant Group SRE・フルスタックエンジニア。SLX の GitOps 実装(ArgoCD + Kubernetes)を担当。(person / sre)
### 2026-06-30 Going from 30 to 30 Million SLOs (SREcon22 EMEA) ingest
- [[Alex Palcuie]](新規) — Google SRE(GCE Compute API チーム・Tech IRT)。per-customer SLO(顧客単位 3,000 万 SLO)と「5 エラーのルール」の設計・実装者。SREcon22 EMEA 発表。(person / sre)
### 2026-06-30 Squish Level Objectives (SREcon20 Americas) ingest
- [[Dave Stanke]](新規) — Google Cloud Platform Developer Advocate。SREcon20 Americas で Customer-centric SRE を論じ、SLO Policy Rationale にユーザー行動データを結び付ける手法と、エラーバジェットの UX 実験活用を提示。(person / sre)
### 2026-06-30 Latency and Availability Error Budgets Done Right at Scale (SREcon20 Americas) ingest
- [[Zendesk]](新規) — カスタマーサービスプラットフォーム企業。Fred Moyer が 1,000 名規模の SLI/SLO/EB 展開を実践。(organization / sre)
- [[Fred Moyer]](更新) — 所属を Zendesk に更新、SREcon20 Americas 発表を追記。複合 SLI と EB 民主化の発表者。
### 2026-06-30 Avoiding Goodhart's Law (SREcon20 Americas) ingest
- [[Marco Coulter]](新規) — AppDynamics(Cisco 傘下)AIOps テクニカルエバンジェリスト。SREcon20 Americas でグッドハートの法則の SRE 応用と 3 次元 SLI/SLO フレームワークを発表。(person / sre)
- [[AppDynamics]](新規) — APM ツール企業(Cisco 傘下)。AIOps・オブザーバビリティ領域。(organization / observability)
### 2026-06-29 SLOs for Data-Intensive Services (SREcon19 EMEA) ingest
- [[Yoann Fouquet]](新規) — Booking.com SRE。データ集約型サービスのデータ品質 SLO 定義実践を SREcon19 EMEA で発表。(person / sre)
- [[Booking.com]](更新) — SREcon19 EMEA でのデータ品質 SLO 発表ソースを追加。
### 2026-06-29 Latency SLOs Done Right (SREcon19 Americas) ingest
- [[Fred Moyer]](新規) — Circonus Developer Evangelist。20 年超のソフトウェア/信頼性経験・Apache Software Foundation メンバー・White Camel Award 2013。SREcon19 Americas でパーセンタイル平均化の誤りとヒストグラム SLO を解説。(person / sre / observability)
- [[Circonus]](更新) — Fred Moyer の所属組織として追記。libcircllhist・circonusllhist・IRONdb の 3 製品情報を拡充。
### 2026-06-29 How Atlassian Is Tackling Error Budgets, Agile Style (SREcon18 Asia) ingest
- [[Gui Vieiro]](新規) — [[Atlassian]] SRE Team Lead。2016 年 Atlassian 参画。Identity チームへのエラーバジェット段階的導入を主導。SREcon18 Asia 登壇。(person / sre)
- [[Atlassian]](新規) — Jira・Confluence 等を開発するソフトウェア企業。マイクロサービス 4 層アーキテクチャ、SRE チームは Observe/Prevent/Improve/Fix の 4 機能。(organization / sre)
### 2026-06-29 SLOs and SLIs in the Real World (SREcon18 Americas) ingest
- [[Matthew Flaming]](新規) — New Relic VP of Site Reliability。SREcon18 Americas でケイパビリティ駆動 SLI/SLO 手法を発表。(person / sre)
- [[Elisa Binette]](新規) — New Relic (@elisabPDX)。Matthew Flaming との共同登壇者。(person / sre)
- [[New Relic]](新規) — オブザーバビリティプラットフォーム企業。SRE 実践事例の発信元。(organization)
### 2026-06-29 Error Budgets and Risks (SREcon15) ingest
- [[Marc Alvidrez]](新規) — Google Senior Staff SRE。2004 年入社。AdSense サービング再実装(2006-2009)を通じてエラーバジェット概念を実践的に発見し、SREcon15 でコミュニティへ初紹介。GFS 初期 SRE チームも経験。(person / sre)
### 2026-06-29 小さくはじめるSLI/SLO / Road to SRE NEXT 2026 神戸 ingest
- [[Narimichi Takamura]](更新) — Road to SRE NEXT 2026 @神戸での SLI/SLO 段階的導入発表を追記。
- [[Topotal]](更新) — 同発表の出典を追記。
### 2026-06-29 AWS Lambda Container Loading / ATC 2023 ingest
- [[Marc Brooker]](新規) — AWS シニアプリンシパルエンジニア。Firecracker MicroVM 共同設計者・Lambda コンテナローディングシステム筆頭著者。分散システム・サーバーレス信頼性の研究者。(person / cloud)
- [[AWS Lambda]](新規) — AWS のサーバーレスイベント駆動型コンピュートサービス(FaaS)。Firecracker MicroVM で隔離。毎秒最大 15,000 コンテナ起動・冷起動 50ms 目標。(product / serverless)
- [[Firecracker]](新規) — AWS のサーバーレス向け軽量仮想化ハイパーバイザー(MicroVM)。KVM ベース・Rust 実装・形式検証済み virtio。NSDI 2020 発表。(product / virtualization)
### 2026-06-29 Raft / ATC 2014 ingest
- [[Diego Ongaro]](新規) — Raft 論文筆頭著者。Stanford University PhD 学生。LogCabin 実装者。(person / distributed)
- [[John Ousterhout]](新規) — Raft 論文共著者。Stanford University 教授。RAMCloud プロジェクト主導者。Tcl 創案者。(person / distributed)
### 2026-06-28 CockroachDB ingest (SIGMOD 2020)
- [[CockroachDB]](新規) — Cockroach Labs が開発する地理分散 SQL DBMS。MVCC + Parallel Commits + HLC で直列化可能分離を汎用ハードウェアで実現。(product / database / distributed / sql)
- [[Cockroach Labs]](新規) — CockroachDB 開発元。SIGMOD 2020 論文には 17 名のエンジニアが名を連ねる。(organization / database)
- [[Rebecca Taft]](新規) — CockroachDB SIGMOD 2020 論文筆頭著者。Cockroach Labs 所属。(person / database / distributed)
- [[Spanner]](更新) — lint-stub から実体ページへ昇格。CockroachDB との詳細比較(一貫性・クロック・レイテンシ)を追記。
### 2026-06-28 F1 ingest (VLDB 2013)
- [[Jeff Shute]](新規) — Google エンジニア。F1 分散 SQL データベース論文筆頭著者。(person / database / distributed)
### 2026-06-28 Amazon MemoryDB ingest (SIGMOD 2024)
- [[Amazon MemoryDB]](新規) — AWS が提供するフルマネージドインメモリ DB サービス。Redis API 完全互換・11 9s 耐久性・4 9s 可用性。2021 年 GA。(product / database / cloud)
- [[Yacine Taleb]](新規) — AWS Canada 所属エンジニア。Amazon MemoryDB 論文筆頭著者。(person / database / cloud)
### 2026-06-28 Amazon Aurora ingest (SIGMOD 2017)
- [[Amazon Aurora (Database)]](新規) — AWS が提供する MySQL/PostgreSQL 互換クラウドネイティブ OLTP DB サービス。2015 年 GA。(product / database / cloud)
- [[Alexandre Verbitski]](新規) — AWS 研究者。Amazon Aurora の筆頭著者。(person / researcher / database)
### 2026-06-28 Characterizing Cloud Computing Hardware Reliability ingest (SoCC 2010)
- [[Kashi Venkatesh Vishwanath]](新規) — Microsoft Research 研究者。データセンターハードウェア信頼性を専門とする。(person / researcher / datacenter)
- [[Nachiappan Nagappan]](新規) — Microsoft Research 研究者。ソフトウェア信頼性工学・実証ソフトウェア工学を専門とする。(person / researcher / reliability)
### 2026-06-28 The Power of Stories ingest (SREcon26 Americas)
- [[Lorin Hochstein]](新規) — Airbnb Staff Software Engineer, Reliability。SRE・インシデント管理・レジリエンスエンジニアリングの実践者。元学者(UNL 准教授、USC/ISI)。(person / sre / incident-management)
- [[Airbnb]](新規) — 民泊プラットフォーム企業。インシデントストーリーテリングの実践組織(「Once Upon an Incident」)。(organization / tech / sre)
### 2026-06-28 Unlock High-Frequency Deployments ingest (SREcon26 Americas)
- [[Ganesh Vernekar]](新規) — Reddit Staff Software Engineer / Prometheus TSDB メンテナー。stale-series compaction の設計者。GitHub: codesome。(person / sre / prometheus)
- [[Reddit]](新規) — 大規模ソーシャルプラットフォーム。Prometheus ベースの大規模オブザーバビリティインフラを運用し、stale-series compaction を本番実験した。(organization / sre)
### 2026-06-28 Reliability Equilibrium ingest (SREcon26 Americas)
- [[Daria Barteneva]](新規) — Microsoft Azure Observability Engineering の Principal SRE。ゲーム理論を SRE に適用する「Reliability Equilibrium」フレームワークの提唱者。SREcon18 EMEA・SREcon26 Americas 登壇。(person / sre / game-theory)
### 2026-06-28 Loop Engineering Working Note (Osmani / HuaShu)
- [[Addy Osmani]](新規) — Google Chrome エンジニア。Loop Engineering 命名者(2026-06-07 ブログ投稿)。Osmani の朝次トリアージループ事例が本ノートの主要実例。(person / software-engineering / agents)
- [[Prithvi Rajasekaran]](新規) — Anthropic エンジニア。長時間動作エージェントアプリ構築中にジェネレータ/エバリュエータ分離パターンの実証知見を得た。(person / agents)
- [[Steve Kaliski]](新規) — Stripe エンジニア。Minions パイプライン(週 1,300+ PR)の設計者。How I AI ポッドキャストで公開。(person / software-engineering)
### 2026-06-28 AI Agents for Incident Investigation ingest (SREcon26 Americas)
- [[Vladyslav Budichenko]](新規) — Vocaly AI 創業者・ソフトウェアエンジニア。SREcon26 Americas で AIエージェントのインシデント調査利活用を実務視点から発表。(person / sre / ai)
- [[Vocaly AI]](新規) — ビジネス向け音声 AIエージェントプラットフォーム。Budichenko 創業。(organization / ai)
### 2026-06-28 Executing Chaos Engineering ingest (SREcon26 Americas)
- [[Leonardo Marques]](新規) — Bradesco SRE Head。25 年超のキャリア。USP デジタルトランスフォーメーション MBA。(person / sre / financial)
- [[Luiz Siqueira]](新規) — Bradesco SRE Manager。SRE MBA・Gremlin Chaos Engineering 認定。IBM/Kyndryl 経験。(person / sre / financial)
- [[Bradesco]](新規) — ブラジル最大級の民間銀行(顧客 7,000 万人以上)。Azure+OpenShift クラウド移行中。カオスエンジニアリング本番適用の実践組織。(organization / financial)
- [[EasyPerform]](新規) — Bradesco 内製カオスエンジニアリングガバナンスプラットフォーム。ガードレール・証跡管理・スケジュール実行・O11y 統合。(product / chaos-engineering)
### 2026-06-28 1年間のポストモーテム運用 ingest (SRE NEXT 2022)
- [[藤原俊一郎]](新規) — 面白法人カヤック SRE。ecspresso・lambroll OSS 作者。ISUCON 11 優勝。ポストモーテム横断統一運用と sre-advisor を主導。(person / sre)
- [[面白法人カヤック]](更新) — SRE NEXT 2022 の登壇ソース追加・Embedded SRE 構造の説明を追記。
### 2026-06-28 Learning from Incidents at Scale ingest (SREcon25 Americas)
- [[Vanessa Huerta Granda]](新規) — Enova テクノロジーマネージャー(レジリエンスエンジニアリング)。元 Jeli.io(Howie ガイド共同執筆)。Resilience in Software Foundation ボードメンバー。(person / sre / incident-management)
- [[Enova]](新規) — フィンテック企業。クロスインシデント分析プログラムを 10 年間運用する実践組織。(organization / fintech / sre)
### 2026-06-28 The Case of the Misnamed Cities ingest (SREcon26 Americas)
- [[Ruben Barroso]](新規) — Google スタッフ SRE。5年以上 STPA/CAST を Google 内部システムに適用。SREcon26 Americas で Google Maps インシデントの CAST 分析を発表。(person / sre / safety-engineering)
- [[Nancy G. Leveson]](新規) — システム安全工学者。CAST・STPA の考案者。"CAST Handbook" 著者。(person / safety-engineering / researcher)
### 2026-06-28 Human Observability of Incident Response ingest (SREcon23 Americas)
- [[Matt Davis]](新規) — FORM.com Site Reliability Architect・音楽家(Craque)。Joint Activity・Common Grounding・Practice of Practice Gamelan を体系化。SREcon23 Americas 2023 登壇者。(person / sre / resilience-engineering)
- [[Pauline Oliveros]](新規) — 作曲家・音楽理論家。Deep Listening 創始者。Davis が The Tuning Meditation を SREcon 体験型導入として使用。(person / music)
- [[Derek Bailey]](新規) — ギタリスト・音楽理論家。「Improvisation has no existence outside of its practice」の著者。Davis がインシデント対応訓練の必要性を論じる際に引用。(person / music)
- [[Laura Maguire]](更新) — Adaptive Choreography(Response Trio: コンダクター・コミュニケーター・問題解決者の三角形)の出典として追加。Managing the Hidden Costs of Coordination(ACM Queue)も本ソースで引用。
### 2026-06-28 Incident Archeology ingest (SREcon23 Americas)
- [[Clint Byrum]](新規) — Spotify スタッフエンジニア・IMOC。SREcon23 Americas で「インシデント考古学」を発表。技術歴 25 年以上。(person / sre / postmortem)
- [[Spotify]](新規) — 音楽・ポッドキャストストリーミング企業。インシデント考古学を 2020〜2021 年の実データで実践した。(organization / sre)
### 2026-06-28 The Repeat Incident Fallacy ingest (SREcon22 EMEA)
- [[Emily Ruppe]](新規) — Jeli.io Solutions Engineer。SREcon22 EMEA で「Repeat Incident Fallacy」を提唱。元 SendGrid・Twilio インシデントコマンドチーム創設メンバー。(person / sre / postmortem)
- [[Laura Maguire]](新規) — レジリエンスエンジニアリング研究者。Jeli.io 所属。Howie: The Post-Incident Guide 共著者。CI/CD = 継続的変化という命題で有名。(person / researcher / resilience-engineering)
### 2026-06-28 A Post Incident Review Review ingest
- [[Tom Partington]](新規) — ANZx SRE(@parmigiana)。SREcon22 APAC でANZx の PIR² プロセスを発表。安全科学(Rasmussen・Dekker・Hollnagel)を PIR 実践に接続。(person / sre / postmortem)
- [[ANZx]](新規) — オーストラリアの金融テクノロジー組織(ANZ グループ傘下)。1000人超・高度規制産業でも根本原因・アクションアイテム・MTTx を除外した PIR を実践。(organization / fintech)
- [[J Paul Reed]] — Chime Staff Incident Operations Manager / 安全科学研究者。PIR 業界調査のデータ提供(SREcon22 APAC)・自動化のアイロニーをAI時代に拡張(SREcon26 Americas)。(person / researcher / postmortem / human-factors)
- [[Chime]] — フィンテック企業。J Paul Reed が Staff Incident Operations Manager として所属。(organization / fintech)
- [[John Allspaw]](新規) — 元 Etsy CTO。Debriefing Facilitation Guide 共著者。ポストモーテム・人的要因研究の第一人者。(person / sre / postmortem)
- [[Jeli]](新規) — インシデント分析ツール企業。Howie: The Post-Incident Guide(2021)発行元。(organization / sre)
- [[Sidney Dekker]](新規) — スウェーデン系安全科学者。Dekker's Tunnel・New View(ヒューマンエラーの新しい見方)の提唱者。Field Guide to Understanding 'Human Error' 等。(person / safety-science)
- [[James Reason]](新規) — 英国認知心理学者。スイスチーズモデル(Swiss Cheese Model)提唱者。(person / safety-science)
- [[Jens Rasmussen]](新規) — デンマーク安全工学者。Workload/Economic/Performance 境界モデルの提唱者。(person / safety-science)
### 2026-06-28 Principled Identification of Root Causes ingest
- [[Laura de Vesine]](新規) — Datadog スタッフエンジニア。SRE・インシデント分析・カオスエンジニアリング専門。SREcon22 EMEA で System/Environment 境界モデルを使った根本原因特定フレームワークを発表。(person / sre / postmortem)
### 2026-06-28 Getting More out of Postmortems (SREcon19Asia) ingest
- [[Ashar Rizqi]](新規) — Blameless Inc. 共同創業者 & CEO。SREcon19 Asia でポストモーテム6要素を体系化。(person / sre)
- [[Blameless]](新規) — シリコンバレー SRE プラットフォームプロバイダー。Ashar Rizqi が共同創業。(organization / sre)
### 2026-06-28 Ditch the Template ingest (SREcon22 EMEA)
- [[Laura Nolan]](新規) — Stanza Systems Principal SWE、元 Google・Slack SRE。SRE Book 共著者。SREcon22 EMEA でナラティブ型 IR 執筆を提唱。(person / sre / postmortem)
- [[Stanza Systems]](新規) — 本番システム制御ソフトウェアのスタートアップ(2022年時点)。Laura Nolan 所属。(organization / sre)
### 2026-06-27 Architecting a Technical Post Mortem ingest
- [[Will Gallego]](新規) — Etsy Staff Systems Engineer。SREcon18 Americas でポストモーテムの哲学と実践を体系化。(person / sre)
### 2026-06-27 Failures and Fixes ingest
- [[Jonathan Sillito]](新規) — Brigham Young University CS 学科。ソフトウェアエンジニアリング実証研究者。30 インシデント定性分析。(person / researcher)
- [[Esdras Kutomi]](新規) — Brigham Young University CS 学科。Sillito との共同研究者。(person / researcher)
- [[Brigham Young University]](新規) — 米国ユタ州プロボの私立大学。Sillito・Kutomi の所属機関。(organization / university)
### 2026-06-27 OTel-Arrow Phase 2 ingest
- [[Apache-Arrow]](新規) — Apache Software Foundation のカラム型インメモリデータフォーマット標準。OTel-Arrow / OTAP のカラム型表現の基盤。(product / data-format)
### 2026-06-27 Do Not Blame Users for Misconfigurations (SOSP'13)
- [[Yuanyuan Zhou]](新規) — UCSD 計算機科学教授。設定ミス研究の責任著者。(person / systems)
- [[Shankar Pasupathy]](新規) — NetApp 研究者。SPEX の産業評価協力者。(person / systems)
- [[NetApp]](新規) — 米国大手ストレージベンダー。SPEX 評価の Storage-A 提供元。(organization / storage)
- [[Tianyin Xu]](更新) — SPEX [SOSP'13] 筆頭著者の記録を追記。
- [[Ding Yuan]](更新) — SOSP'13 共著者として UCSD 在籍時の研究を追記。
### 2026-06-27 障害箇所特定・根本原因分析 11 論文一括 ingest
- [[BSODiag]](新規) — クラウドインフラ障害のバッチサーバー障害診断フレームワーク。(product / cloud-infra / rca)
- [[COCA]](新規) — コード知識で強化した生成的 RCA フレームワーク。(product / rca / llm)
- [[RADICE]](新規) — 因果サブグラフ出力型 RCA。(product / causal / rca)
- [[LogInsight]](新規) — LLM ベースのログ障害診断システム。(product / log / llm)
- [[FaaSRCA]](新規) — サーバーレスアプリケーション向け RCA。(product / serverless / rca)
- [[FL-AIer]](新規) — DéjàVu 拡張の障害箇所特定システム。(product / fault-localization)
- [[UniDiag]](新規) — TKG ベースの統合マイクロサービス障害診断。(product / microservice / knowledge-graph)
- [[LasRCA]](新規) — LLM + 小型分類器のワンショット RCA。(product / rca / llm)
- [[DiagFusion]](新規) — マルチモーダル障害診断フレームワーク。(product / rca)
- [[Christopher Lohse]](新規) — IBM Research Europe。(person / causal)
- [[Tao Duan]](新規) — Xi'an Jiaotong University / BSODiag 筆頭著者。(person / rca)
- [[Pinghui Wang]](新規) — Xi'an Jiaotong University。(person / rca)
- [[Andrea Tonon]](新規) — Huawei Ireland / RADICE 筆頭著者。(person / causal / rca)
- [[He Jiang]](新規) — Dalian Maritime University。(person / fault-localization)
- [[Shikai Guo]](新規) — FL-AIer 共著者。(person / fault-localization)
- [[Shiyu Ma]](新規) — Nankai University / LogInsight 共著者。(person / log)
- [[Tong Xiao]](新規) — Tsinghua University / LogInsight 共著者。(person / log)
- [[Rui Ren (Alibaba)]](新規) — Alibaba DAMO / SLIM 筆頭著者。(person / fault-localization)
- [[Jingbang Yang]](新規) — Alibaba DAMO / SLIM 共著者。(person / fault-localization)
- [[Linxiao Yang]](新規) — Alibaba DAMO。(person / fault-localization)
- [[Xinyue Gu]](新規) — Alibaba DAMO。(person / fault-localization)
- [[Liang Sun (Alibaba DAMO)]](新規) — Alibaba DAMO Academy。(person / fault-localization)
- [[Yongqi Han]](新規) — Tongji University / LasRCA 筆頭著者。(person / rca)
- [[Qingfeng Du]](新規) — Tongji University / LasRCA 共著者。(person / rca)
- [[Yongxin Zhao]](新規) — Nankai University / UniDiag 共著者。(person / microservice)
- [[Junzhou Zhao]](新規) — Xi'an Jiaotong University。(person / rca)
- [[Huawei Ireland Research Center]](新規) — Huawei の欧州研究拠点。(organization)
- [[Dalian Maritime University]](新規) — 大連海事大学。(organization)
- [[Di-Matrix]](新規) — LasRCA 産業パートナー。(organization)
- [[Tongji University]](新規) — 同済大学。(organization)
- [[Dan Pei]](更新) — DéjàVu・UniDiag・LogInsight 追記。(person / aiops)
- [[Shenglin Zhang]](更新) — UniDiag・LogInsight 追記。(person / aiops)
- [[Yongqian Sun]](更新) — LogInsight・UniDiag 追記。(person / aiops)
- [[Zeyan Li]](更新) — DéjàVu 筆頭著者として明確化。(person / fault-localization)
- [[Michael R. Lyu]](更新) — COCA 追記。(person / rca)
- [[Guangba Yu]](更新) — COCA・FaaSRCA 追記。(person / rca)
- [[Pengfei Chen]](更新) — FaaSRCA 追記。(person / rca)
- [[Alibaba Cloud]](更新) — BSODiag 研究追記。(organization)
- [[IBM Research]](更新) — 因果発見マイクロサービス研究追記。(organization)
- [[Xi'an Jiaotong University]](更新) — BSODiag 研究追記。(organization)
- [[Train-Ticket]](更新) — FL-AIer での利用事例追記。(dataset)
- [[RCACopilot]](更新) — COCA との比較追記。(product / rca)
- [[Nengwen Zhao]](更新) — DéjàVu 共著追記。(person / aiops)
- [[ZTE Corporation]](更新) — LogInsight 参加追記。(organization)
- [[China Mobile Research]](更新) — LogInsight 参加追記。(organization)
### 2026-06-27 RCA・障害箇所特定・集合通信診断 9 論文一括 ingest
- [[LLMRCA]](新規) — LLM アプリケーション特化の多段 RCA フレームワーク。マルチモーダルオブザーバビリティデータ統合。(product / aiops / rca)
- [[Gou Tan]](新規) — LLMRCA 筆頭著者。(person / aiops)
- [[MetaRCA]](新規) — メタ因果知識による汎化可能 RCA フレームワーク。(product / aiops / rca)
- [[BiAn]](新規) — 本番規模ネットワークの LLM ベース障害箇所特定システム。(product / networking / fault-localization)
- [[Guyue Liu]](新規) — BiAn 共著者。(person / networking)
- [[TWIST]](新規) — 分布内介入による因果推論ベース RCA 手法。(product / causal / rca)
- [[eARCO]](新規) — プロンプト最適化による効率的自動 RCA。(product / aiops / rca)
- [[PromptWizard]](新規) — プロンプト最適化フレームワーク。eARCO が活用。(product / llm / prompt-optimization)
- [[GALA]](新規) — グラフ拡張 LLM エージェントワークフロー RCA。(product / aiops / rca / agent)
- [[RCAEval]](新規) — RCA 手法評価フレームワーク。(product / aiops / benchmark)
- [[Ennan Zhai]](更新) — BiAn (SIGCOMM 2025) 共著者として追記。(person / networking)
- [[Drishti Goel]](新規) — Robust Root Cause Diagnosis 共著者。(person / causal)
- [[Lokesh Nagalapatti]](新規) — Robust Root Cause Diagnosis 共著者。(person / causal)
- [[Amit Sharma]](新規) — Microsoft Research India。因果推論研究者。(person / causal)
- [[Sunita Sarawagi]](新規) — IIT Bombay 教授。(person / machine-learning)
- [[IIT Bombay]](新規) — インド工科大学ボンベイ校。(organization)
- [[Giuliano Casale]](新規) — Imperial College London。(person / systems)
- [[Imperial College London]](新規) — ロンドンの工科大学。(organization)
- [[Hans-Arno Jacobsen]](新規) — University of Toronto 教授。(person / systems)
- [[Shuai Liang]](新規) — CCL-D 関連。(person / hpc)
- [[Yifang Tian]](新規) — CCL-D 関連。(person / hpc)
- [[Ant Group]](更新) — BALANCE / KPIRoot+ 関連で追記。(organization)
- [[University of Chinese Academy of Sciences]](更新) — 論文関連で追記。(organization)
- [[China Unicom Software Research Institute]](新規) — 中国聯通軟件研究院。(organization)
- [[PetShop]](新規) — マイクロサービス RCA ベンチマーク環境。(product / benchmark)
### 2026-06-27 データベースノブチューニング・自律 DB 3 論文
- [[Jiale Lao]](新規) — Sichuan University。GPTuner(VLDB 2024)の筆頭著者。(person / database)
- [[Mingjie Tang]](新規) — Sichuan University 准教授。GPTuner(VLDB 2024)の責任著者。(person / database)
- [[OtterTune]](新規) — CMU 発の ML ベース DBMS 自動チューニングシステム。SIGMOD 2017 で発表。(product / database)
- [[Dana Van Aken]](新規) — [[Carnegie Mellon University]]。OtterTune(SIGMOD 2017)の筆頭著者。(person / database)
- [[Andrew Pavlo]](新規) — [[Carnegie Mellon University]] データベースグループ。OtterTune の共著者。(person / database)
- [[openGauss]](新規) — Huawei 発のオープンソース自律データベース。(product / database)
- [[Guoliang Li]] / [[Xuanhe Zhou]] / [[Tsinghua University]] — openGauss(VLDB 2021)の参照を追記。(entity 更新)
- [[Cornell University]](新規) — Immanuel Trummer の所属。DB-BERT 関連。(organization)
### 2026-06-27 データベース異常診断・RCA 8 論文
- [[Lingsen Yan]](新規) — [[Huazhong University of Science and Technology]]。SDN(ICDE 2025)の筆頭著者。(person / database)
- [[Bolong Zheng]](新規) — [[Huazhong University of Science and Technology]]。SDN(ICDE 2025)の責任著者。(person / database)
- [[Xiaofang Zhou]](新規) — HKUST。SDN(ICDE 2025)共著者。(person / database)
- [[Huazhong University of Science and Technology]](新規) — 中国の研究大学。(organization)
- [[Biao Ouyang]](新規) — [[East China Normal University]]。RCRank(VLDB 2025)の筆頭著者。(person / database)
- [[Chenjuan Guo]](新規) — [[East China Normal University]]。RCRank 共著者。(person / database)
- [[East China Normal University]](新規) — 上海の研究大学。DASE がデータ科学の拠点。(organization)
- [[Aalborg University]](新規) — デンマーク。Christian S. Jensen の所属。(organization)
- [[Vikramank Singh]](新規) — [[Amazon Web Services]]。Vista(Amazon Science)の筆頭著者。(person / database)
- [[Lizhi Liao]](新規) — [[University of Waterloo]]。FSE 2023 産業経験報告の筆頭著者。(person / software-engineering)
- [[Heng Li]](新規) — Polytechnique Montréal。FSE 2023 共著者。(person / software-engineering)
- [[Weiyi Shang]](新規) — [[University of Waterloo]]。FSE 2023 共著者。(person / software-engineering)
- [[Chaoyu Chen]](新規) — [[Ant Group]]。BALANCE(PACMMOD 2023)の共同筆頭著者。(person / rca)
- [[OceanBase]](新規) — Ant Group 発のクラウドネイティブ分散 DB。(product)
- [[Hanzhang Wang]](新規) — [[eBay]]。GRANO(VLDB 2019)の筆頭著者。(person / rca)
- [[eBay]](新規) — EC プラットフォーム。NuData 分散 DB を自社運用。(organization)
- [[NuData]](新規) — eBay 内製の地理分散データベース。GRANO の対象システム。(product)
- [[Vimalkumar Jeyakumar]](新規) — [[Cisco Tetration Analytics]]。ExplainIt!(SIGMOD 2019)の筆頭著者。(person / rca)
- [[Navindra Yadav]](新規) — [[Cisco Tetration Analytics]]。ExplainIt! 共著者。(person / rca)
- [[Cisco Tetration Analytics]](新規) — Cisco のデータセンター監視・分析プラットフォーム。(product)
- [[Tim Kraska]](更新) — Vista(Amazon RDS)の共著者として追記。(person / database)
- [[Jianguo Li]](更新) — BALANCE(PACMMOD 2023)の責任著者として追記。(person / rca)
- [[Ant Group]](更新) — BALANCE の開発組織として追記。(organization)
- [[Amazon Web Services]](更新) — Vista の開発組織として追記。(organization)
### 2026-06-27 データベースノブチューニングサーベイ 2 本 ingest
- [[Xinyang Zhao]](新規) — [[Tsinghua University]]。ノブチューニングサーベイ(TKDE 2023)の共同筆頭著者。(person / database)
- [[Limeng Zhang]](新規) — [[University of Adelaide]] CREST。クラウド DB チューニングサーベイ(arXiv 2024)の筆頭著者。(person / database)
- [[M. Ali Babar]](新規) — [[University of Adelaide]] CREST ディレクター。クラウド DB チューニングサーベイの共著者。(person / software-engineering / database)
- [[University of Adelaide]](新規) — オーストラリアの研究大学。CREST がクラウド DB チューニング研究の拠点。(organization)
- [[Guoliang Li]](更新) — ノブチューニングサーベイ(TKDE 2023)の責任著者として追記。(person / database)
- [[Xuanhe Zhou]](更新) — ノブチューニングサーベイ(TKDE 2023)の共同筆頭著者として追記。(person / database)
- [[Tsinghua University]](更新) — ノブチューニングサーベイ(TKDE 2023)の所属組織として追記。(organization)
### 2026-06-27 DB-BERT SIGMOD 2022 論文 ingest
- [[Immanuel Trummer]](新規) — Cornell University データベース研究者。NLP と強化学習を DB チューニングに融合する研究の推進者。DB-BERT の提案者。(person / database / nlp)
- [[Cornell University]](更新) — Immanuel Trummer の所属機関として追記。DB-BERT 論文との関連を記録。(organization / university)
### 2026-06-26 SRE NEXT 2022 AIOps研究録 スライド ingest
- [[TSifter]](新規) — マイクロサービス性能異常の原因診断に向け、異種メトリクスを事前指定せず因果グラフ生成前に削減する時系列前処理手法。(product / aiops / time-series)
- [[坪内佑樹]] / [[Meltria]](更新) — 講演の著者情報と、故障注入・負荷生成・運用データ蓄積を行うデータセット生成ワークフローを追記。(person / dataset / aiops)
### 2026-06-26 DEMi 分散実行最小化 (NSDI 2016) 論文 ingest
- [[Colin Scott]](新規) — UC Berkeley 博士課程(当時)。分散システムデバッギング研究者。DEMi および STS (SIGCOMM 2014) の筆頭著者。(person / distributed-systems / debugging)
- [[Scott Shenker]](新規) — UC Berkeley 教授 / ICSI。分散システム・ネットワーキングの権威。DEMi の共著者。(person / distributed-systems / networking)
- [[George Necula]](新規) — UC Berkeley 教授。プログラミング言語・形式検証の研究者。DEMi の共著者。(person / distributed-systems / programming-languages)
- [[Arvind Krishnamurthy]](更新) — DEMi (NSDI 2016) の共著者として追記。(person / distributed-systems)
- [[Aurojit Panda]](更新) — DEMi (NSDI 2016) の共著者として追記。(person / distributed-systems)
### 2026-06-26 ソフトウェア信頼性工学 2 論文 ingest
- [[James J. Cusick]](新規) — ITSM プロセス管理ディレクター。AT&T / Bell Labs / Dell 経験。SRE 歴史研究者・実践者。(person / software-reliability)
- [[John Musa]](新規) — (1933–2009) Bell Labs の SRE パイオニア。Musa-Okumoto SRGM 開発者。"Software Reliability Engineering" の用語造語者。(person / software-reliability)
- [[Norman F. Schneidewind]](新規) — (1928–2015) Naval Postgraduate School 名誉教授、IEEE Fellow。Schneidewind 信頼性モデル開発者。NASA スペースシャトル信頼性予測に貢献。(person / software-reliability)
- [[Martin L. Shooman]](新規) — MIT / Polytechnic Institute of New York。"Probabilistic Reliability" (1968) 著者。ソフトウェア信頼性の統計的基盤確立者。(person / software-reliability)
- [[ISSRE]](新規) — IEEE International Symposium on Software Reliability Engineering。1990 年 Michael R. Lyu が創設。SRE の研究と実務を結ぶ主要学会。(organization / software-reliability)
### 2026-06-26 SREcon23 EMEA スライド ingest(From Sysadmins to Flying Unicorns)
- [[Guillaume Hérail]](新規) — [[Sony Interactive Entertainment]] Staff SRE。SREcon23 EMEA 登壇者。(person / sre)
- [[Gilberto Müller]](新規) — [[Sony Interactive Entertainment]] SRE Manager。SREcon23 EMEA 登壇者。(person / sre)
- [[Sony Interactive Entertainment]](新規) — PlayStation Cloud Gaming SRE。SRE 組織変革事例の組織。(organization / sre)
### 2026-06-26 HDD: Hierarchical Delta Debugging 論文 ingest
- [[Ghassan Misherghi]](新規) — UC Davis 大学院生(ICSE 2006 時点)。HDD の筆頭著者。(person / software-engineering / debugging)
- [[Zhendong Su]](新規) — UC Davis 教授。プログラミング言語・ソフトウェアテスティング研究者。HDD の共著者。(person / software-engineering / programming-languages)
### 2026-06-26 データベース/分散システム異常診断 6 論文一括 ingest
- [[iSQUAD]](新規) — Alibaba OLTP Database 向け間欠的遅延クエリ診断フレームワーク。TOPIC クラスタリングとベイズ事例モデルで根本原因を自動診断。(product / aiops / database)
- [[Tim Kraska]](新規) — [[MIT CSAIL]] 教授。OSprey・学習済みインデックスなどデータシステム × ML の研究者。(person / database / ml)
- [[OSprey]](新規) — MIT CSAIL の OS メトリクス事前学習トランスフォーマー。因子分解アーキテクチャでクエリレイテンシ予測のシステム間汎化を実現。(product / database / ml)
- [[Apache IoTDB]](新規) — Apache Software Foundation の分散時系列データベース。MultiLog・LogDB・RBAD の評価プラットフォーム。(product / database / time-series)
- [[RBAD]](新規) — Peking University の Raft ログベース異常診断手法。ランタイム収集オーバーヘッドほぼゼロで分散ストレージの異常を診断。(product / aiops / distributed-storage)
- [[ZTE Corporation]](更新) — DBPA ベンチマーク共著組織として追記。(organization / telecom)
- [[Minghua Ma]](更新) — iSQUAD 第一著者(PVLDB 2020)として追記。(person / aiops)
- [[Dan Pei]](更新) — iSQUAD 共著者として追記。(person / aiops)
- [[Shenglin Zhang]](更新) — iSQUAD 共著者として追記。(person / aiops)
- [[Samuel Madden]](更新) — OSprey 共著者(MIT CSAIL)として追記。(person / database)
- [[MIT CSAIL]](更新) — OSprey 研究チームの所属として追記。(organization)
- [[Lingzhe Zhang]](更新) — MultiLog・LogDB・RBAD の筆頭著者として追記。(person / aiops)
- [[Tong Jia]](更新) — MultiLog・LogDB・RBAD の責任著者として追記。(person / aiops)
- [[Ying Li]](更新) — MultiLog・LogDB・RBAD の責任著者として追記。(person / aiops)
- [[Peking University]](更新) — MultiLog・LogDB・RBAD・DBPA の主要所属として追記。(organization)
- [[Bin Cui]](更新) — DBPA 共著者として追記。(person / database)
- [[Shiyue Huang]](更新) — DBPA 第一著者として追記。(person / database)
### 2026-06-26 SREcon22 APAC 動画 ingest (Reliability Map)
- [[Aaron Bowden]](新規) — Google Cloud Professional Services SRE Practice Lead JAPAC。[[Reliability Map (r9y.dev)]] オープンソースプロジェクト主導者。(person / sre / google)
### 2026-06-26 SONiC Workshop Japan 2026 スライド ingest
- [[海老澤健太郎]](新規) — [[Arrcus]] プリンシパルエンジニア。SONiC・Ultra Ethernet・Scale-Up ネットワーキング専門家。『実践 SONiC 入門』著者。(person / networking)
- [[Arrcus]](新規) — ネットワーク OS ベンダー。(organization / networking)
### 2026-06-26 ReCycle (SOSP 2024) 論文 ingest
- [[Swapnil Gandhi]](新規) — [[Stanford University]]。ReCycle 筆頭著者。耐障害分散訓練(person / distributed / fault-tolerant-training)
- [[Christos Kozyrakis]](新規) — [[Stanford University]] 教授。ReCycle 共著者。コンピュータアーキテクチャ・システム(person / systems / architecture)
- [[ReCycle]](新規) — Stanford University の耐障害 DNN 訓練システム。データ並列冗長性 + パイプラインバブル活用でスペアサーバ不要(product / fault-tolerant-training)
### 2026-06-26 Alibaba HPN (SIGCOMM 2024) 論文 ingest
- [[Alibaba HPN]](更新) — SIGCOMM 2024 論文の詳細を追記。2 層デュアルプレーン + 非スタック型デュアル ToR + レール最適化(product / networking)
- [[Kun Qian]](更新) — Alibaba HPN 筆頭著者として追記(person / networking)
- [[Alibaba Cloud]](更新) — HPN 開発元として追記(organization)
### 2026-06-26 なめらかなシステム (DICOMO2018) 論文 ingest
- [[栗林健太郎]](新規) — [[GMOペパボ]] ペパボ研究所。「なめらかなシステム」提唱者(第一著者、DICOMO2018)。(person / systems-thinking)
- [[三宅悠介]](更新) — DICOMO2018 第二著者として追記。
- [[Ryosuke Matsumoto]](更新) — DICOMO2018 第三著者(当時 [[GMOペパボ]])として追記。
- [[GMOペパボ]](更新) — なめらかなシステム論文(DICOMO2018)の所属組織として追記。
---
### 2026-06-26 SC24 GPU-to-GPU Communication 論文 ingest
- [[Tiziano De Matteis]](新規) — [[Vrije Universiteit Amsterdam]]。HPC ネットワーキング研究者。SC 2024 共著者。(person / hpc)
- [[Zebin Ren]](新規) — [[Vrije Universiteit Amsterdam]]。HPC ネットワーキング研究者。NWO 助成(OCENW.KLEIN.561)。SC 2024 共著者。(person / hpc)
- [[Animesh Trivedi]](新規) — [[IBM Research]] Europe(旧: [[Vrije Universiteit Amsterdam]])。システム・ストレージ・HPC ネットワーキング研究者。SC 2024 共著者。(person / systems)
- [[Duncan Roweth]](新規) — HPE Cray。HPC インターコネクト・Slingshot 技術者。SC 2024 共著者。(person / hpc / networking)
- [[Daniele De Sensi]](更新) — SC 2024 GPU 間インターコネクト包括評価研究(第一著者)を追記。
- [[Lorenzo Pichetti]](更新) — SC 2024 共著者として追記。
- [[Flavio Vella]](更新) — SC 2024 共著者として追記。
- [[Torsten Hoefler]](更新) — SC 2024 共著者(ETH Zürich 代表)として追記。
---
### 2026-06-26 arXiv 2401.00134 Unicron 論文 ingest
- [[Tao He (Alibaba)]](新規) — [[Alibaba Group]]。Unicron の筆頭著者。(person / machine-learning)
- [[Jingren Zhou]](新規) — [[Alibaba Group]]。Unicron の共著者(最終著者)。(person / machine-learning)
- [[Unicron]](新規) — Alibaba Group 開発の LLM 訓練向け自己修復ワークロードマネージャ。Megatron 基盤。(product / distributed / machine-learning)
- [[Kun Qian]](更新) — Unicron 共著者として追記。
- [[Alibaba Group]](更新) — Unicron の開発・評価を追記。
---
### 2026-06-26 HotNets 2024 I've Got 99 Problems But FLOPS Ain't One 論文 ingest
- [[Costin Raiciu]](新規) — [[University Politehnica of Bucharest]] / [[Broadcom]]。マルチパストランスポート・データセンターネットワーキング研究者。HotNets 2024 対応著者。(person / networking)
- [[University Politehnica of Bucharest]](新規) — ルーマニア工科大学(UPB)。[[Costin Raiciu]] のネットワーキング研究グループが拠点。(organization / networking)
- [[Broadcom]](更新) — [[Costin Raiciu]] の兼任所属として HotNets 2024 論文を追記。
---
### 2026-06-26 ICPADS 2024 Generic and ML Workloads in an HPC Datacenter 論文 ingest
- [[Xiaoyu Chu]](新規) — [[Vrije Universiteit Amsterdam]]。ICPADS 2024 等貢献筆頭著者。HPC データセンターのワークロード特性化研究者。(person / hpc)
- [[Alexandru Iosup]](新規) — [[Vrije Universiteit Amsterdam]] 教授・`atlarge-research` 主宰。HPC・分散システムのワークロード特性化・スケジューリング研究。(person / hpc / distributed-systems)
- [[SURF]](新規) — オランダの IT インフラ協同組合。SURF Lisa(338 ノード HPC)の運営組織で本研究のデータ提供者。(organization / hpc)
- [[Ivona Brandic]](新規) — [[TU Wien]] 教授。クラウドコンピューティング・HPC・エネルギー効率研究者。(person / hpc / cloud-computing)
- [[Vrije Universiteit Amsterdam]](更新) — `atlarge-research` グループと HPC ワークロード特性化研究を追記。
### 2026-06-26 ICSE 2023 Quality Issues of DL Platform 論文 ingest
- [[Yanjie Gao]](更新) — ICSE 2023 品質問題実証研究(筆頭著者)を追記。
- [[Hongyu Zhang]](更新) — ICSE 2023 共著者(Chongqing University)として追記。
- [[Microsoft Research]](更新) — Platform-X 品質問題 ICSE 2023 研究を追記。
---
### 2026-06-26 Demystifying NCCL 論文 ingest
- [[ATLAHS]](新規) — ETH Zürich SPCL のアプリケーション追跡駆動型ネットワークシミュレーションツールチェーン。NCCL 通信パターンを誤差 5% 未満で再現。
- [[Siyuan Shen]](新規) — ETH Zürich SPCL。Demystifying NCCL の等貢献筆頭著者・ATLAHS 共著者。
- [[Zhiyi Hu]](新規) — ETH Zürich SPCL。Demystifying NCCL の等貢献筆頭著者。
- [[NCCL]](更新) — 体系的内部解析論文を追記。
- [[Torsten Hoefler]](更新) — Demystifying NCCL を追記。
### 2026-06-28 SREcon26 Americas TrainCheck 講演 ingest
- [[Ryan Huang]](新規) — [[University of Michigan]] [[OrderLab]]。SREcon 2026 Americas にて [[Yuxuan Jiang]] と共同発表。
- [[Yuxuan Jiang]](更新) — SREcon26 Americas 登壇ソース追記。
- [[TrainCheck]](更新) — SREcon26 Americas 講演ソース追記。
### 2026-06-26 OSDI 2025 TrainCheck 論文 ingest
- [[Yuxuan Jiang]](新規) — [[University of Michigan]] [[OrderLab]]。TrainCheck 筆頭著者。
- [[Peng Huang]](新規) — [[University of Michigan]] [[OrderLab]] PI。DL システム信頼性・不変条件推論研究。
- [[TrainCheck]](新規) — DL 訓練サイレントエラー検知フレームワーク(OSDI 2025)。
- [[OrderLab]](新規) — [[University of Michigan]] 内研究室。[[Peng Huang]] が PI。
- [[University of Michigan]](更新) — TrainCheck/OrderLab の所属機関として追記。
### 2026-06-24 ClickHouse PVLDB 2024
- [[ClickHouse]](新規) — オープンソースカラム型 OLAP データベース。
- [[ClickHouse Inc|ClickHouse Inc.]](新規) — ClickHouse の開発・維持企業。
- [[Robert Schulze]](新規) — ClickHouse Inc. エンジニア、論文筆頭著者。
- [[Alexey Milovidov]](新規) — ClickHouse 創設者。
### 2026-06-24 SREcon スライド 7 件一括取り込み (anomaly detection / monitoring)
- [[Yu Chen (Baidu)]](更新) — SREcon19 Asia でゴールデンシグナルの異常検知を発表。
- [[Andrew Clegg]](新規) — [[Etsy]] 主席データサイエンティスト。時系列マイニング(SAX/DTW/類似度検索)を発表。(person / data-science)
- [[Etsy]](新規) — EC プラットフォーム。Kale/Skyline パイプラインで時系列異常検知を実運用。(organization / e-commerce)
- [[Ivan Shubin]](新規) — [[Booking.com]] シニア SRE。Granomaly(統計ベース異常検知サービス)を開発。(person / sre)
- [[Booking.com]](新規) — オンライン旅行プラットフォーム。AI/ML なしの統計ベース異常検知を実運用。(organization / travel)
- [[Ian Neidel]](新規) — [[Netflix]] ソフトウェアエンジニア。ゲーム QoE メトリクスの変化点検知を発表。(person / sre)
- [[Netflix]](更新) — ゲームストリーミングの異常検知を追記。
- [[Open Connect]](新規) — Netflix のグローバル CDN。ゲームストリーミングの品質モニタリング基盤。(product / cdn)
- [[Joseph Cirella]](新規) — [[Bloomberg]] シニアエンジニア。PELT 変化点検知で性能レグレッション自動検知を構築。(person / sre)
- [[Shanthini Velan]](新規) — [[Bloomberg]] エンジニア。CI/CD パイプラインでの変化点検知を発表。(person / sre)
- [[Zhaogang Wang]](更新) — [[Alibaba Group]] SRE。ビジネストレンド異常検知の SREcon17 Asia 発表を追記。
- [[Alibaba Group]](更新) — ビジネストレンド異常検知の産業事例として追記。
- [[Xianping Qu]](新規) — [[Baidu]] エンジニア。4 モジュール自動異常検知プラットフォームを発表。(person / sre)
- [[Baidu]](更新) — SREcon15 での自動異常検知プラットフォームを追記。
### 2026-06-23 SREcon18 Americas Automatic Metric Screening
- [[Yu Chen (Baidu)]](更新) — SREcon18 Americas でサービス診断向け自動メトリクススクリーニングを発表。SREcon17 Asia のアラート疲労対策と ISSRE 2019 FluxRank の間に位置する Baidu SRE/AIOps 実践として追記。
- [[Baidu]](更新) — サービス診断向けメトリクススクリーニングの発表組織として追記。
### 2026-06-23 SREcon17 Americas Practical Monitoring and Alerting
- [[Jamie Wilkinson]](更新) — [[Google]] SRE。SREcon17 Americas で時系列・分布・Prometheus を使った監視とアラート設計を説明し、SREcon18 Asia の SLO バーンレートアラート体系化へつながる前段を示した。(person / sre / prometheus)
- [[Prometheus]](更新) — ラベル付き時系列、記録ルール、トポロジ集約を大規模監視設定の抽象として使う実践例を追記。
### 2026-06-23 SREcon16 Europe Alerting for Distributed Systems
- [[Björn Rabenstein]](新規) — [[SoundCloud]] Production Engineer。[[Prometheus]] 主要開発者の一人。元 [[Google]] SRE。SREcon16 Europe で症状ベースアラーティングと時系列ベース監視を発表。(person / sre / prometheus)
- [[SoundCloud]](新規) — 音声共有サービス企業。発表では Prometheus を含む監視・アラーティング改善の事例対象。(organization / monitoring)
- [[Prometheus]](更新) — Rabenstein SREcon16 Europe の時系列ベースディスク満杯アラートと alerts 用語の文脈を追記。
### 2026-06-23 SREcon17 Asia Draining the Flood スライド
- [[Argus (Baidu)]](新規) — Baidu の内製監視システム。数百の分散サービスを監視し、4 施策導入で 85% のアラート削減を達成。(product / monitoring)
### 2026-06-23 SREcon17 Europe Over-Monitoring and Alert Fatigue スライド
- [[Kishore Jalleda]](新規) — Yahoo プロダクションエンジニアリング Sr. Director(元 [[Zynga]] SRE 責任者)。Clean Room イニシアティブでアラートバジェットによるインセンティブ設計を主導。(person / sre)
- [[Zynga]](新規) — ソーシャルゲーム開発・運営企業。Clean Room イニシアティブの実施組織。偽アラーム 90% 削減を達成。(organization / gaming)
### 2026-06-23 SREcon21 Spike Detection スライド
- [[Nishant Singh]](新規) — [[LinkedIn]] シニア SRE(Production-SRE チーム)。修正 Z スコアによるアラート相関のスパイク分離を実装・発表。(person / sre)
### 2026-06-23 SREcon19 EMEA Adaptive Paging スライド
- [[Luis Mineiro]](新規) — [[Zalando SE]] SRE 責任者。分散トレーシングの因果関係を活用したアラートルーティング手法 [[Adaptive Paging]] を提案。(person / sre)
- [[Zalando SE]](新規) — ヨーロッパ最大級のオンラインファッションプラットフォーム。マイクロサービスアーキテクチャで [[Adaptive Paging]] を実運用。(organization / e-commerce)
### 2026-06-23 SRE NEXT 2023 Warning アラート自動調査スライド
- [[池田将士]](新規) — [[面白法人カヤック]] その他事業部 SRE チームのデータエンジニア。[[prepalert]] と shimesaba の作者。(person / sre / data-engineering)
- [[面白法人カヤック]](新規) — Web サービス・ゲーム・広告などを手がける日本の企業。発表では SRE チームの Warning アラート運用改善の舞台。(organization / web-service)
- [[prepalert]](新規) — Mackerel webhook を起点に CloudWatch Logs Insights / S3 Select / Redshift Data API / plugin providers から補助情報を集め、アラートメモへ貼るツール。(repository / alert-management)
- [[Mackerel]] / [[SRE NEXT]](更新) — Warning アラート運用改善と SRE NEXT 2023 発表として追記。
### 2026-06-23 ハイブリッドアテンション論文(Rethinking Hybrid Architectures)
- [[Zhiyuan Liu]](新規) — Tsinghua University。NLP・LLM。(person / machine-learning)
- [[Xu Han]](新規) — Tsinghua University。LLM。(person / machine-learning)
- [[Chaojun Xiao]](新規) — Tsinghua University。効率的アテンション・長コンテキスト。(person / machine-learning)
- [[OpenBMB]](新規) — 清華大学 NLP グループ連携オープンソースコミュニティ。(organization)
- [[Tsinghua University]](更新) — Rethinking Hybrid Architectures ソースを追加。
### 2026-06-23 LLM 基盤論文 4 本一括(InstructGPT / Chinchilla / Sparsely-Gated MoE / ReAct)
- [[Long Ouyang]](新規) — OpenAI。InstructGPT 第一著者。(person / machine-learning / alignment)
- [[Jordan Hoffmann]](新規) — DeepMind。Chinchilla 第一著者。(person / machine-learning / scaling)
- [[DeepMind]](新規) — Google 傘下の AI 研究所。Chinchilla・AlphaGo 等。(organization / machine-learning)
- [[Azalia Mirhoseini]](新規) — Google Brain → Google DeepMind。Sparsely-Gated MoE 共著者。(person / machine-learning)
- [[Geoffrey Hinton]](新規) — ニューラルネットワーク・深層学習の先駆者。(person / machine-learning)
- [[Jeffrey Dean]](新規) — Google Senior Fellow。MapReduce・TensorFlow 等の中核貢献。(person / systems / machine-learning)
- [[Quoc V. Le]](新規) — Google Brain。Sparsely-Gated MoE 共著者。(person / machine-learning)
- [[Shunyu Yao]](新規) — Princeton University。ReAct 第一著者。(person / machine-learning / agents)
- [[Karthik Narasimhan]](新規) — Princeton University。ReAct 共著者。(person / machine-learning)
- [[Princeton University]](新規) — 米国プリンストン大学。(organization)
- [[Jared Kaplan]](新規) — Johns Hopkins University / Anthropic 共同創業者。Neural Scaling Laws 著者。(person / machine-learning / scaling)
- [[Google Brain]](更新) — Sparsely-Gated MoE・Chinchilla 等の参照を追加。(organization / machine-learning)
- [[OpenAI]](更新) — InstructGPT の参照を追加。(organization / machine-learning)
- [[Noam Shazeer]](更新) — Sparsely-Gated MoE 第一著者としての参照を追加。(person / machine-learning)
### 2026-06-21 マイクロサービス RCA・マルチモーダル障害診断 7 論文一括(LocaleXpert / UniTok / MRCA / HolisticRCA / Medicine / ChangeLLM / DeepHunt)
- [[Zhouruixing Zhu]](新規) — LocaleXpert 共著者。(person / aiops)
- [[Yunhao Zhang]](新規) — [[Shanghai Jiao Tong University]]。UniTok 第一著者。(person / time-series)
- [[Ruiying Qi]](新規) — [[Shanghai Jiao Tong University]]。UniTok 共著者。(person)
- [[Jiale Zheng]](新規) — [[Huawei Noah's Ark Lab]]。UniTok 共著者。(person)
- [[Jianfeng Zhang (Huawei)]](新規) — [[Huawei Noah's Ark Lab]]。UniTok 共著者。(person)
- [[Lujia Pan]](新規) — [[Huawei Noah's Ark Lab]]。UniTok 共著者。(person)
- [[Junchi Yan]](新規) — [[Shanghai Jiao Tong University]]。UniTok 責任著者。(person)
- [[Huawei Noah's Ark Lab]](新規) — Huawei の AI 研究部門。UniTok の共同研究機関。(organization)
- [[Yidan Wang]](新規) — MRCA 第一著者。(person / aiops)
- [[Yongqi Han]](新規) — [[Tongji University]]。HolisticRCA 第一著者。(person / aiops)
- [[Qingfeng Du]](新規) — [[Di-Matrix]]。HolisticRCA 共著者。(person / aiops)
- [[Tongji University]](新規) — 上海の総合大学。HolisticRCA の主要所属。(organization)
- [[Di-Matrix]](新規) — AIOps プラットフォーム企業。HolisticRCA の産業連携先。(organization / product)
- [[Lei Tao]](新規) — Medicine 第一著者。(person / aiops)
- [[Zhengdan Li]](新規) — Medicine 共著者。(person / aiops)
- [[Yuchi Ma]](新規) — ChangeLLM/SCELM 第一著者。(person / aiops)
- [[Qiuai Fu]](新規) — ChangeLLM/SCELM 共著者。(person / aiops)
- [[Chetan Bansal]](更新) — [[Microsoft Research]]。ChangeLLM 共著者として追記。(person / aiops)
- [[Minghua Ma]](更新) — Medicine 共著者として追記。(person / aiops)
- [[Shenglin Zhang]](更新) — Medicine・DeepHunt 共著者として追記。(person / aiops)
- [[Dan Pei]](更新) — Medicine・DeepHunt 共著者として追記。(person / aiops)
- [[Pinjia He]](更新) — ChangeLLM 共著者として追記。(person / aiops)
- [[Nankai University]](更新) — DeepHunt の主要所属として追記。(organization)
### 2026-06-20 マイクロサービス RCA 6 論文一括(TraceRank / LogCluster / LogKG / FSF / Nezha / Eadro)
- [[Zicheng Huang]] — [[Sun Yat-sen University]]。TraceRank 共著者。(person / aiops)
- [[Jian-Guang Lou]] — [[Microsoft Research]]。LogCluster 共著者。(person / aiops / log-analysis)
- [[Xuewei Chen]] — [[Microsoft Research]]。LogCluster 共著者。(person)
- [[Yu Zhang]] — [[Microsoft Corporation]]。LogCluster 共著者。(person)
- [[Yicheng Sui]] — [[Tsinghua University]]。LogKG 第一著者。(person / aiops / log-analysis)
- [[Jesus Rios]] — [[IBM Research]]。FSF(Faulty Service Finder) 共著者。(person / aiops)
- [[Larisa Shwartz]] — [[IBM Research]]。FSF 共著者。(person / aiops)
- [[Cheryl Lee]] — [[The Chinese University of Hong Kong]]。Eadro 第一著者。(person / aiops)
- [[Tianyi Yang]] — [[The Chinese University of Hong Kong]]。Eadro 共著者。(person / aiops)
- [[Yuxin Su]] — [[Sun Yat-sen University]]。Eadro 共著者。(person / aiops)
- [[Zibin Zheng]] — [[Sun Yat-sen University]]。Nezha 共著者。(person / aiops)
### 2026-06-20 Energy statistics (JSPI 2013)
- [[Gábor J. Székely]] — NSF / ハンガリー科学アカデミー Rényi 数学研究所。エネルギー統計・距離共分散の提唱者。(person / statistics)
- [[Maria L. Rizzo]] — Bowling Green State University。エネルギー統計・距離共分散の共同提唱者。R `energy` パッケージ開発者。(person / statistics)
### 2026-06-20 Odin (NSDI 2018)
- [[Matt Calder]] — [[Microsoft]] / [[University of Southern California|USC]]。Odin 筆頭著者。CDN 計測・エニキャスト研究者。(person / networking)
- [[Ethan Katz-Bassett]] — [[Columbia University]] 教授。Odin 共著者。(person / networking)
- [[Ganesh Ananthanarayanan]] — [[Microsoft]]。Odin 共著者。(person / systems)
- [[Ratul Mahajan]] — [[Intentionet]](元 [[Microsoft Research]])。Odin 共著者。(person / networking)
- [[Columbia University]] — ニューヨーク市の私立研究大学。(organization)
- [[Intentionet]] — ネットワーク管理スタートアップ。(organization)
### 2026-06-20 分散トレーシング基礎論文 5 本一括
- [[Mike Y. Chen]] — [[UC Berkeley ROC Project]]。Pinpoint 第一著者。(person / distributed-systems)
- [[Emre Kıcıman]] — [[Stanford University]]。Pinpoint 共著者。(person / distributed-systems)
- [[Armando Fox]] — [[Stanford University]] → UC Berkeley。Pinpoint 共著者、ROC プロジェクト主導。(person / distributed-systems)
- [[Eric Brewer]] — [[UC Berkeley ROC Project]]。Pinpoint 共著者、CAP 定理提唱者。(person / distributed-systems)
- [[Pinpoint]] — J2EE 計装 + 統計的クラスタリングによる障害コンポーネント自動特定システム(DSN 2002)。(product / fault-localization)
- [[Paul Barham]] — [[Microsoft Research]] Cambridge。Magpie 第一著者。(person / systems)
- [[Rebecca Isaacs]] — [[Microsoft Research]] Cambridge。Magpie 共著者。(person / systems)
- [[Richard Mortier]] — [[Microsoft Research]] Cambridge → Cambridge University。Magpie 共著者。(person / systems)
- [[Magpie]] — イベントベースリクエスト抽出とワークロードモデリングシステム(HotOS IX 2003)。(product / distributed-tracing)
- [[Xu Zhao]] — [[University of Toronto]]。lprof 第一著者。(person / systems)
- [[Ding Yuan]] — [[University of Toronto]]。lprof 共著者。(person / systems)
- [[Michael Stumm]] — [[University of Toronto]]。lprof 共著者。(person / systems)
- [[lprof]] — 非侵入リクエストフロープロファイラ(OSDI 2014)。(product / distributed-tracing)
- [[Jonathan Mace]] — [[Brown University]]。Pivot Tracing / Canopy 著者。(person / distributed-tracing)
- [[Ryan Roelke]] — [[Brown University]]。Pivot Tracing 共著者。(person / distributed-tracing)
- [[Rodrigo Fonseca]] — [[Brown University]]。Pivot Tracing 共著者、X-Trace 著者。(person / distributed-tracing)
- [[Pivot Tracing]] — 動的因果モニタリングフレームワーク(SOSP 2015)。(product / distributed-tracing)
- [[Jonathan Kaldor]] — [[Facebook]]。Canopy 第一著者。(person / distributed-tracing)
- [[Scuba]] — Facebook のインメモリ分析データベース。Canopy のバックエンド。(product / analytics)
- [[Brown University]] — 米国ロードアイランド州。Pivot Tracing・Canopy の研究拠点。(organization / university)
- [[UC Berkeley ROC Project]] — Recovery-Oriented Computing プロジェクト。Pinpoint の研究拠点。(organization / research-project)
### 2026-06-20 ISSTA 2016 — Practitioners' Expectations on FL
- [[Xin Xia]] — [[Zhejiang University]]。FL 採用期待調査共著者。(person / software-engineering)
- [[Pavneet Singh Kochhar]] — [[Singapore Management University]]。FL 採用期待調査第一著者。(person / software-engineering)
- [[Shanping Li]] — [[Zhejiang University]]。FL 採用期待調査共著者。(person / software-engineering)
- [[Singapore Management University]] — シンガポールの私立大学。SMU。(organization / university)
- [[David Lo]] — 追記: [[@2016__ISSTA__Practitioners' Expectations on Automated Fault Localization]] の共著者。
### 2026-06-20 IWQoS 2020 — MicroCause
- [[Yuan Meng]] — [[Tsinghua University]] / BNRist。MicroCause 第一著者。(person / aiops / rca)
- [[Ruru Zhang]] — [[Nankai University]]。MicroCause 共著者。(person / aiops)
- [[Zhilong Hu]] — [[Nankai University]]。MicroCause 共著者。(person / aiops)
- [[Yiyin Zhang]] — [[Alibaba Group]]。MicroCause 共著者(データ提供)。(person / aiops)
- [[Chenyang Jia]] — [[Alibaba Group]]。MicroCause 共著者。(person / aiops)
- [[Zhaogang Wang]] — [[Alibaba Group]]。MicroCause 共著者。(person / aiops)
### 2026-06-20 arXiv — A Tutorial on Kernel Density Estimation and Recent Advances
- [[Yen-Chi Chen]] — [[University of Washington]] 統計学部。KDE チュートリアルの著者。(person / nonparametric-statistics)
### 2026-06-20 JMLR — DirectLiNGAM
- [[Shohei Shimizu]] — [[Osaka University]] 産業科学研究所。LiNGAM・DirectLiNGAM 提案者。(person / causal-discovery)
- [[Aapo Hyvärinen]] — [[University of Helsinki]]。独立成分分析(ICA)の主要研究者。(person / ICA / causal-discovery)
- [[Kenneth Bollen]] — University of North Carolina。構造方程式モデル(SEM)の権威。(person / structural-equation-models)
- [[Osaka University]] — 大阪大学。産業科学研究所が LiNGAM 系研究の中心。(organization)
- [[University of Helsinki]] — ヘルシンキ大学。ICA 研究の拠点。(organization)
### 2026-06-19 Signal Processing — Selective review of offline change point detection methods
- [[Charles Truong]] — CMLA, CNRS, [[ENS Paris-Saclay]]。変化点検知サーベイの筆頭著者、[[ruptures]] 開発者。(person / change-point-detection / signal-processing)
- [[Laurent Oudre]] — L2TI, University Paris 13。変化点検知サーベイ共著者。生体信号処理。(person / signal-processing)
- [[Nicolas Vayatis]] — CMLA, CNRS, [[ENS Paris-Saclay]]。変化点検知サーベイ共著者。(person / machine-learning)
- [[ruptures]] — 多変量時系列のオフライン変化点検知 Python ライブラリ。BSD ライセンス。13 種のコスト関数と 5 種の探索手法をモジュール式に組み合わせ可能。(repository / change-point-detection)
### 2026-06-19 PVLDB — Time-Series Clustering: A Comprehensive Study
- [[John Paparrizos]] — [[The Ohio State University]] / Aristotle University of Thessaloniki。k-Shape 原著者。84手法の包括的ベンチマークを主導。(person / time-series / clustering)
- [[UCR Time Series Archive]] — 128データセットを収録する時系列分類・クラスタリングの標準ベンチマーク。MOMENT のデータ汚染問題が指摘された。(dataset / time-series / benchmark)
- [[The Ohio State University]] — 既存ページ更新。John Paparrizos と時系列クラスタリング研究を追記。(organization)
### 2026-06-19 Boris Tane Blog — The Software Development Lifecycle Is Dead
- [[Boris Tane]] — ソフトウェアエンジニア・テックブロガー。AI エージェント時代における SDLC の解体と[[コンテキストエンジニアリング]]の台頭を論じた。(person / software-development / ai-native)
### 2026-06-19 CSUR — D'ya Like DAGs? A Survey on Structure Learning and Causal Discovery
- [[Matthew J. Vowels]] — CVSSP, [[University of Surrey]]。構造学習・因果発見包括サーベイの筆頭著者。(person / causal-discovery / structure-learning)
- [[Necati Cihan Camgoz]] — CVSSP, [[University of Surrey]]。サーベイ共著者。(person)
- [[Richard Bowden]] — CVSSP, [[University of Surrey]]。サーベイ共著者。(person)
- [[University of Surrey]] — 既存ページ更新。CVSSP・因果発見研究を追記。(organization)
### 2026-06-19 Frontiers in Genetics — Review of Causal Discovery Methods
- [[Clark Glymour]] — [[Carnegie Mellon University]] 哲学科。因果発見アルゴリズムの先駆者。*Causation, Prediction, and Search* 共著者。(person / causal-discovery)
- [[Kun Zhang]] — CMU 哲学科。PNL 因果モデル、カーネルベース独立性検定の研究者。本レビューの責任著者。(person / causal-discovery)
- [[Peter Spirtes]] — CMU 哲学科。PC・FCI アルゴリズムの共同開発者。TETRAD 中核開発者。(person / causal-discovery)
- [[Carnegie Mellon University]] — 既存ページ更新。因果発見アルゴリズムの発祥拠点を追記。(organization)
### 2026-06-19 Physics Reports — Signal propagation in complex networks
- [[Peng Ji]] — [[Fudan University]] ISTBI 所属、責任著者。複雑ネットワーク・非線形動力学・脳インスパイア知能の研究者。(person / complex-networks)
- [[Jürgen Kurths]] — Potsdam Institute for Climate Impact Research / Humboldt University の非線形動力学・複雑系の世界的権威。(person / complex-networks / nonlinear-dynamics)
- [[Matjaž Perc]] — [[University of Maribor]] の物理学者。複雑ネットワーク・進化ゲーム理論・社会ダイナミクス専門。(person / complex-networks)
- [[University of Maribor]] — スロベニアの研究大学。Perc の所属機関。(organization / university)
### 2026-06-19 CSUR — Anomaly Detection: A Survey
- [[Varun Chandola]] — ACM Computing Surveys 2009 の異常検知サーベイ第一著者。点/文脈/集合異常と 6 技法群の taxonomy を整理した。(person / anomaly-detection)
- [[Arindam Banerjee]] — Chandola+ 2009 異常検知サーベイ共著者。既存 [[Subho S. Banerjee]] とは別人物。(person / anomaly-detection)
- [[Vipin Kumar]] — Chandola+ 2009 異常検知サーベイ共著者。(person / anomaly-detection)
- [[University of Minnesota]] — Chandola+ 2009 の著者所属機関。(organization / university)
### 2026-06-19 SREcon19 EMEA — Latency SLOs Done Right
- [[Heinrich Hartmann]] — SRE・テレメトリ分析・統計の実務家。SREcon19 EMEA でレイテンシ SLO におけるパーセンタイル集約不能性と、ログ・カウンタ・ヒストグラムの実装経路を説明。(person / sre / telemetry)
- [[Circonus]] — モニタリングおよびテレメトリ分析に関わる企業。Hartmann の SREcon19 EMEA 資料ではヒストグラムメトリクス、IronDB、`libcircllhist` の文脈で登場。(organization / observability)
### 2026-06-19 GRLIA — Graph-based Incident Aggregation
- [[GRLIA]] — グラフ表現学習ベースのオンラインインシデント集約フレームワーク。Huawei Cloud Networking サービスで評価・展開された。(product / aiops / incident-management)
- [[OpsPAI]] — GRLIA の公開実装 `https://github.com/OpsPAI/grlia` を置く GitHub 名前空間。(repository / aiops)
- [[Xuemin Wen]] / [[Xiao Ling]] — Huawei 所属の実務者。GRLIA 共著者。(person / aiops)
- [[Zhuangbin Chen]] / [[Jinyang Liu]] / [[Yuxin Su]] / [[Hongyu Zhang]] / [[Yongqiang Yang]] / [[Michael R. Lyu]](更新) — GRLIA 論文との関連を追記。
- [[Huawei Cloud]] / [[The Chinese University of Hong Kong]] / [[University of Newcastle]](更新) — GRLIA の本番データ・共同研究関係を追記。
### 2026-06-18 SpeakerDeck — AIスーパーコンピュータにおけるLLM学習処理性能の計測と可観測性
- [[Yuuki Tsubouchi]](更新) — 情報処理学会中国支部主催講演会で、[[SAKURAONE]] の GPT-3 175B 学習ベンチマーク、ジョブ履歴分析、OTel + Grafana によるリソース分析、GPU ゼロコード計装、RoCE 能動プロービングを AI スパコン可観測性の統合視点として説明(person / ai-supercomputer / observability)
- [[SAKURA Internet]] / [[SAKURAONE]](更新) — SAKURAONE の 3 クラスタ構成、MLPerf Training 型ベンチマーク、可観測性基盤を講演資料から追記(organization / product / hpc)
- [[R-Pingmesh]](更新) — `yuuki/rpingmesh` 実装を、RoCE ネットワーク問題とアプリケーション問題の切り分けを狙う運用実装として追記(product / rdma / monitoring)
### 2026-06-18 Contextual Retrieval — Anthropic
- [[Daniel Ford]](新規) — [[Anthropic]] エンジニア。Contextual Retrieval の提案者(2024-09-19 Anthropic Engineering Blog)。(person / rag / information-retrieval)
- [[Anthropic]](更新) — Contextual Retrieval の一次資料([[@2024__Anthropic Engineering Blog__Introducing Contextual Retrieval]])を出典に追加し、失敗率の数値を一次資料(5.7% → 1.9%)に修正。
### 2026-06-18 PyTorch Conference 2025 — LMCache + NIXL
- [[Moein Khazraee]] — [[NVIDIA]] の Senior Architect。[[NIXL]] と [[LMCache]] の連携による KV キャッシュ転送を PyTorch Conference 2025 で発表(person / llm-serving / kv-cache)
- [[Junchen Jiang]](更新) — [[LMCache]] と [[NIXL]] による長コンテキスト推論向け KV キャッシュ転送/保存設計の発表を追記(person / llm-serving / kv-cache)
- [[LMCache]] / [[NIXL]](更新) — LMCache + NIXL 資料から、メモリ登録、メタデータ交換、UCX/GDS/OBJ バックエンド、ストレージ設定例を追記(product / kv-cache / transfer)
### 2026-06-18 KV キャッシュ・GPU クラスタ論文 5 本
- [[Xingda Wei]] / [[Jinbo Han]] — SJTU IPADS。KVCache Cache in the Wild 著者(person / kv-cache / workload-characterization)
- [[Qizhen Weng]] — [[Hong Kong University of Science and Technology]] / [[Alibaba Group]]。MLaaS in the Wild 筆頭著者(person / gpu-cluster / scheduling)
- [[Alibaba PAI]] — Alibaba の ML-as-a-Service プラットフォーム(product / mlaaas)
- [[Alibaba GPU Cluster Trace]] — Alibaba PAI の公開 GPU クラスタトレース(dataset / gpu-cluster)
- [[Jiayi Yao]] / [[Junchen Jiang]] — [[University of Chicago]]。CacheBlend 著者(person / kv-cache / rag)
- [[CacheBlend]] — 非プリフィックス KV キャッシュ再利用の選択的再計算システム(product / llm-serving)
- [[Huan Yang]] — [[Central South University]]。KVShare 筆頭著者(person / kv-cache / multi-tenant)
- [[KVShare]] — マルチテナント KV キャッシュ再利用システム(product / llm-serving)
- [[Central South University]] — 中南大学(organization)
- [[Yucheng Li]] / [[Huiqiang Jiang]] — Microsoft / University of Surrey。SCBench 著者(person / kv-cache / benchmark)
- [[SCBench]] — KV キャッシュ中心の長コンテキスト手法ベンチマーク(product / benchmark)
- [[University of Chicago]] / [[University of Surrey]] — 組織(organization)
### 2026-06-18 FlashAttention シリーズ 4 本 + AIBrix ingest
- [[Tri Dao]] — [[Stanford University]] → [[Princeton University]] / [[Together AI]]。FlashAttention シリーズ(FA1-FA4)の第一著者。IO-aware 厳密アテンションの創始者(person / attention / gpu-optimization)
- [[Jay Shah]] — [[Colfax Research]]。FlashAttention-3/4 の共同第一著者。Hopper/Blackwell 向けカーネル設計(person / attention / gpu-optimization)
- [[AIBrix]](更新) — 一次論文 [[@2025__arXiv__AIBrix...]] を反映。description を協調設計哲学・機能一覧に拡充(product / llm-serving / cloud-native)
- [[Together AI]](更新) — [[Tri Dao]] の所属と FlashAttention シリーズとの関係を追加(organization / llm)
### 2026-06-18 From Attention to Disaggregation 充実化に伴う追加
- [[NVIDIA Dynamo]] — NVIDIA フルスタックと密結合した分離型推論サービングフレームワーク。Planner / Smart Router(KV-aware)/ KV Cache Block Manager / Prefill-Decode Workers / [[NIXL]] を統合。671B 級モデル本番運用を想定(product / llm-serving / pd-disaggregation)。**既存 [[Dynamo]]([[Amazon]] の KVS)とは別物**として独立ページに切り出し
- [[AIBrix]] — Kubernetes 上の LLM 推論クラウドネイティブ制御プレーン。LLM-specific Autoscaler、GPU Optimizer、Unified AI Runtime、Distributed KV Cache でマルチテナント・コスト最適化(product / llm-serving / cloud-native)
### 2026-06-18 Mooncake — KVCache-centric Disaggregated Architecture for LLM Serving
- [[Ruoyu Qin]] / [[Zheming Li]] / [[Weiran He]] — [[Moonshot AI]]。Mooncake 論文の主著者(person / llm-serving / moonshot)
- [[Mingxing Zhang]] / [[Yongwei Wu]] / [[Weimin Zheng]] — [[Tsinghua University]] MadSys グループ。Mooncake の共同設計者(person / llm-serving / tsinghua)
- [[Xinran Xu]] — [[Moonshot AI]]。Mooncake 論文の対応著者(person / llm-serving / moonshot)
- [[Mooncake]](更新) — Kimi の本番サービングプラットフォームとして詳細を更新。3 プール分離・Conductor・CPP・Layer-wise Prefill・過負荷指向スケジューリングを追記(product / llm-serving / kv-cache / pd-disaggregation)
### 2026-06-18 MPLS JAPAN 2025 — KV cache sharing with IOWN APN
- [[田仲顕至]] — [[NTT]] デバイスイノベーションセンタの研究者。MPLS JAPAN 2025 で [[IOWN APN]] を用いた KV キャッシュ共有による LLM 推論高速化を発表(person / llm-serving / kv-cache)
- [[NTT]] — IOWN APN と KV キャッシュ共有による分散 LLM 推論構想を発表した組織(organization / telecom / llm-serving)
- [[IOWN APN]] — 低遅延・広帯域の All-Photonics Network。小規模データセンター間の KV キャッシュ共有ネットワークとして扱われる(product / network / iown)
### 2026-06-18 LLM 推論 KV キャッシュ管理/分離型推論 6 論文
- [[SGLang]] — 構造化された言語モデルプログラムを frontend/runtime 協調で実行するシステム。RadixAttention、圧縮 FSM、API speculative execution を備える(product / llm-serving / kv-cache)
- [[P-D-Serve]] — Huawei の Ascend/MindSpore 上で数万 NPU に商用展開された Prefill-Decode 分離型 LLM サービングシステム(product / llm-serving / pd-disaggregation)
- [[Woosuk Kwon]] / [[Yuhan Liu]] / [[Srinivasa Rao Aravilli]] / [[Yibo Jin]] / [[Zixuan Zhou]] — 今回取り込んだ PagedAttention、LMCache、From Attention to Disaggregation、P/D-Serve、Efficient Inference Survey の第一著者(person / llm-serving)
- [[Tensormesh Inc]] / [[Infinigence-AI]] / [[Capital One]] — 今回取り込んだ LMCache、Efficient Inference Survey、From Attention to Disaggregation の著者所属組織(organization / llm-serving)
- [[vLLM]] / [[LMCache]](更新) — PagedAttention 原典、LMCache 論文、SGLang/P-D-Serve との関係を追記(product / kv-cache)
### 2026-06-18 LLM 推論サービング論文 2 本
- [[DistServe]] — [[Peking University]] / [[UC San Diego]] / [[StepFun]] の OSDI 2024 LLM サービングシステム。Prefill と Decode を分離し per-GPU Goodput を最大化(product / llm-serving)
- [[Yinmin Zhong]] / [[Shengyu Liu]] / [[Junda Chen]] / [[Jianbo Hu]] / [[Xuanzhe Liu]] — [[DistServe]] 論文の著者(person / llm-serving)
- [[Ranran Zhen]] / [[Juntao Li]] / [[Yixin Ji]] / [[Zhenlin Yang]] / [[Tong Liu]] / [[Min Zhang]] — INLG 2025「Taming the Titans」サーベイの [[Soochow University]] 側著者(person / llm-serving)
- [[Qingrong Xia]] / [[Xinyu Duan]] / [[Zhefeng Wang]] / [[Baoxing Huai]] — INLG 2025「Taming the Titans」サーベイの [[Huawei Cloud]] 側著者(person / llm-serving)
- [[Soochow University]] / [[UC San Diego]] / [[StepFun]] — 今回の LLM 推論サービング論文で追加した所属機関(organization / university / llm-serving)
- [[Peking University]] / [[Huawei Cloud]] / [[Yibo Zhu]] / [[Xin Jin]] / [[Hao Zhang]] / [[vLLM]](更新) — DistServe と INLG 2025 サーベイからの所属・共著・ベースライン評価を追記(entity / llm-serving)
### 2026-06-18 SpeakerDeck — 推論基盤のパフォーマンス検証と最適化戦略
- [[道下幹也]](更新) — 第 3 回 vLLM roundup Community Meetup Tokyo で SLO/SLA ベースの推論基盤最適化、PD Disaggregation、Mooncake Store による KV Cache Reuse/Sharing の実測を発表(person / llm-serving / gpu)
- [[SAKURA Internet]] / [[高火力 PHY]](更新) — 高火力 PHY 上の LLM 推論基盤検証として PD 分離と KV Cache Reuse/Sharing の実測を追加(organization / product / gpu)
- [[vLLM]] / [[LMCache]] / [[Mooncake]](更新) — `vllm bench serve`、LMCache、Mooncake Store を使った KV Cache Reuse/Sharing 検証を追加(product / llm-serving / kv-cache)
### 2026-06-17 自動化のアイロニー後続 2 論文(Baxter+ ECCE2012 / Strauch IEEE-THMS2017)
- [[Gordon Baxter]] — [[University of St Andrews]]、ヒューマンファクター研究者(person / human-factors)
- [[John Rooksby]] — [[University of St Andrews]]、コンピュータサイエンス研究者(person / human-factors)
- [[Barry Strauch]] — [[National Transportation Safety Board]] 退職、事故調査官・ヒューマンファクター専門家(person / accident-investigation)
- [[University of St Andrews]] — スコットランドの大学、EPSRC LSCITS プロジェクト拠点(organization / university)
- [[National Transportation Safety Board]] — 米国の独立安全調査機関(organization / safety)
### 2026-06-17 マイクロサービスベンチマーク/データセット 4 論文一括
- [[Christina Delimitrou]] — [[Cornell University]]、DeathStarBench PI(person / microservices)
- [[Yu Gan]] — [[Cornell University]]、DeathStarBench 第一著者(person / microservices)
- [[Cornell University]] — DeathStarBench 開発元(organization / university)
- [[Davide Taibi]] — [[University of Oulu]] M3S Cloud Group、OSS-MS dataset 主導(person / microservices)
- [[Tomas Cerny]] — [[Baylor University]]→[[University of Arizona]]、microservice test 研究(person / microservices)
- [[University of Oulu]] — M3S Cloud Group、OSS-MS dataset 拠点(organization / university)
- [[Baylor University]] — Cerny + Smith ほか、microservice test benchmark(organization / university)
- [[eShopOnContainers]] — Microsoft .NET の e-commerce microservice 参照アプリ(repository / microservices)
- [[EvoMaster]] — REST API 自動テスト生成ツール(product / testing)
- [[World of Code]] — FLOSS 173M projects のリポジトリマイニング基盤(dataset / mining-software-repositories)
- [[Software Competence Center Hagenberg]] — オーストリアの応用 SW 研究センター、TrainTicketTrace 著者(organization / research-center)
- [[Pirmin Urbanke]] — SCCH Hagenberg、TrainTicketTrace 主著者(person / microservices)
- [[Stefan Fischer]] — SCCH Hagenberg、TrainTicketTrace 共著者(person / microservices)
- [[Dario Amoroso d'Aragona]] — [[Tampere University]]、OSS-MS dataset 主著者(person / microservices)
- [[Alexander Bakhtin]] — [[University of Oulu]]、OSS-MS dataset + LO2 dataset(person / microservices)
- [[Tampere University]] — Amoroso d'Aragona の所属(organization / university)
### 2026-06-17 アラート管理論文 3 本(Zha+ / VOCE / SkyNet)
- [[Junjie Zha]] — [[State Grid Jiangsu Electric Power]] 所属、Zha+ Electronics2024 の第一著者(person / aiops)
- [[Xinwen Shan]] — [[State Grid Jiangsu Electric Power]] 所属、Zha+ Electronics2024 共著者(person / aiops)
- [[Jiaxin Lu]] — [[State Grid Jiangsu Electric Power]] 所属、Zha+ Electronics2024 共著者(person / aiops)
- [[Jiajia Zhu]] — [[State Grid Jiangsu Electric Power]] 所属、Zha+ Electronics2024 Supervisor(person / aiops)
- [[Zihan Liu]] — [[State Grid Jiangsu Electric Power]] 所属、Zha+ Electronics2024 共著者(person / aiops)
- [[State Grid Jiangsu Electric Power]] — 中国・江蘇省の州営電力会社。Information & Telecommunication Branch が IT 運用 100K アラート × 130 ストームの本番データを提供(organization / utility / aiops)
- [[Jia Chen (Fudan)]] — [[Fudan University]] 所属、VOCE 第一著者。OAS(ICSE2022)も(person / aiops)
- [[Xiaolei Chen]] — [[Fudan University]] 所属、VOCE 共著者(person / aiops)
- [[Jie Shi]] — [[Fudan University]] 所属、VOCE 共著者(person / aiops)
- [[Peng Wang (Fudan)]] — [[Fudan University]] グループリーダー、VOCE corresponding author(person / aiops)
- [[Wei Wang (Fudan)]] — [[Fudan University]] 所属、VOCE 共著者(person / aiops)
- [[Bo Yang]] — [[Alibaba Cloud]] 所属、SkyNet co-primary author。preprocessor + hierarchical alert tree 設計(person / aiops / networking)
- [[Huanwu Hu]] — [[Alibaba Cloud]] 所属、SkyNet co-primary author。NetAssistant 等にも参画(person / aiops / networking)
- [[Yifan Li]] — [[Alibaba Cloud]] 所属、SkyNet co-primary author(person / aiops / networking)
- [[Tao Lin (Alibaba)]] — [[Alibaba Cloud]] 所属、SkyNet corresponding author([[Ennan Zhai]] と共に)(person / aiops / networking)
- [[node2vec]] — Grover & Leskovec(KDD 2016)のグラフ表現学習。Zha+ 2024 で enterprise topology graph 埋め込みに使用(product / graph-embedding)
- [[Sentence-BERT]] — Reimers & Gurevych(EMNLP 2019)の文埋め込みモデル。Zha+ 2024 でアラートテキスト類似度計算に使用(product / nlp)
- [[FT-tree]] — Zhang+(IWQoS 2017)の syslog template 構築アルゴリズム。SkyNet で Syslog 千種類規模を type 分類するために採用(product / log-parsing / network)
- [[Eigenvector Centrality]] — グラフ理論のノード中心性指標。VOCE が fault propagation graph 上の originating source 推定に採用(product / graph-theory)
## Person
- [[Mark Burgess]] — CFEngine 開発者。"Computer Immunology" (USENIX LISA 1998) でシステム管理の工学化を主張し、約束理論(promise theory)を提唱した先駆者(person / sre / systems-administration)
- [[Ryota Yoshikawa]] — [[Topotal]] CTO。`@rrreeeyyy` として SpeakerDeck で SRE/AI 運用関連資料を公開し、"Reliability in the Age of AI" で AI 時代の開発速度と信頼性の再調整を論じた(person / sre / aiops)
- [[佐藤竜馬]] — [[National Institute of Informatics]] 助教。LLM の注意ヘッド分類・機構的解釈性・[[グラフニューラルネットワーク]]・[[モデルパラメータ算術]]・最適輸送の研究者。著書 3 冊。ブログ「ジョイジョイジョイ」著者。joisino 系 13 記事(2024-09〜2026-03)を 2026-06-16 にバッチ ingest 済み(person / machine-learning / llm)
- [[Zeyuan Allen-Zhu]] — [[Meta FAIR]] 所属。[[Yuanzhi Li]] と共同で [[Physics of Language Models]] シリーズを主導し、合成データ+線形プロービングで LLM の[[知識操作]]・[[知識容量スケーリング則]](パラメータ 1 つにつき約 2 ビット記憶)・[[文脈自由文法]]学習などの普遍則を抽出する制御実験を確立した([[joisino-言語モデルの物理学-2025]])(person / machine-learning / llm / interpretability)
- [[Yuanzhi Li]] — Mohamed bin Zayed University of Artificial Intelligence(MBZUAI)所属。[[Zeyuan Allen-Zhu]] と共同で [[Physics of Language Models]] シリーズを主導([[joisino-言語モデルの物理学-2025]])(person / machine-learning / llm / interpretability)
- [[Yann LeCun]] — [[Meta FAIR]] Chief AI Scientist・NYU 教授。次トークン予測のみによる学習が[[LLM意味表象]]の人間整合性を高めず、世界モデルや[[認知意味論]]に学んだ自己教師あり表現学習を提唱([[joisino-LLMと言葉の感じ方-2026]])(person / machine-learning / ai-research)
- [[Michelle Brush]] — [[Google]] Engineering Director, SRE。Google Compute Engine と Persistent Disk の信頼性を担う。SREcon26 Americas 2026 で、AI エージェント時代の複雑システム信頼性、汎用緩和、実験、リスク先行開発を論じた。(person / sre / reliability)
- [[David DeWitt]] — [[University of Wisconsin]]-Madison Computer Sciences Department。[[Gamma]] プロジェクトの設計者。シェアードナッシング並列データベース・データパーティショニング・ハッシュ結合の研究リーダー([[@1992__CACM__Parallel Database Systems The Future of High Performance Database Systems]])(person / database / parallel)
- [[Abdelouahab Khelifati]] — [[University of Fribourg]] / [[eXascaleInfolab]]。TSM-Bench([[@2023__PVLDB__TSM-Bench - Benchmarking Time Series Database Systems for Monitoring Applications]])の筆頭著者。TS-LSH データ生成手法の提案者。(person / database / time-series)
- [[Mourad Khayati]] — [[University of Fribourg]] / [[eXascaleInfolab]]。TSM-Bench 共著者。時系列欠損値回復(ORBITS・Mind the Gap)の研究者。(person / database / time-series)
- [[Djellel Difallah]] — [[NYU Abu Dhabi]]。TSM-Bench 共著者。データベースベンチマーク・クラウドソーシング研究者。(person / database)
- [[Philippe Cudré-Mauroux]] — [[University of Fribourg]] 教授・[[eXascaleInfolab]] 主宰。TSM-Bench 上級著者。エクサスケールデータ管理・時系列処理の研究リーダー。(person / database / time-series)
- [[Bryan Cantrill]] — [[Sun Microsystems]] Solaris Kernel Development。[[DTrace]] の筆頭発明者。USENIX ATC 2004([[@2004__USENIX-ATC__Dynamic Instrumentation of Production Systems]])の第一著者(person / observability / systems)
- [[Michael Shapiro]] — [[Sun Microsystems]] Solaris Kernel Development。[[DTrace]] 共同発明者。USENIX ATC 2004 の第二著者(person / observability / systems)
- [[Adam Leventhal]] — [[Sun Microsystems]] Solaris Kernel Development。[[DTrace]] 共同発明者。USENIX ATC 2004 の第三著者(person / observability / systems)
- [[Steven McCanne]] — [[LBNL]](Lawrence Berkeley National Laboratory)所属(1993 年時点)。[[BPF]](BSD Packet Filter)の共同発明者。[[@1993__USENIX__The BSD Packet Filter A New Architecture for User-level Packet Capture]] の第一著者(person / networking / observability / systems)
- [[Van Jacobson]] — [[LBNL]](Lawrence Berkeley National Laboratory)所属(1993 年時点)。TCP 輻輳制御・ヘッダ圧縮の発明者としても著名。[[BPF]] の共同発明者。[[@1993__USENIX__The BSD Packet Filter A New Architecture for User-level Packet Capture]] の共著者(person / networking / systems)
- [[Phillip Wenig]] — [[Hasso Plattner Institute]] / Philipps University of Marburg。時系列異常検知の包括的評価([[@2022__PVLDB__Anomaly Detection in Time Series - A Comprehensive Evaluation]])の筆頭著者。[[TimeEval]]・[[GutenTAG]] の開発者(person / anomaly-detection / time-series)
- [[Sebastian Schmidl]] — [[Hasso Plattner Institute]] / University of Potsdam。同包括的評価の第一著者。[[TimeEval]] フレームワークの共同開発者(person / anomaly-detection / time-series)
- [[Abdul Fatir Ansari]] — [[AWS AI Labs]]([[Amazon Web Services]])所属。Chronos([[@2024__arXiv__Chronos Learning the Language of Time Series]]、TMLR 2024)・[[Chronos-2]] の equal contribution 筆頭著者(* 印)。時系列トークナイゼーション(スケーリング+均一量子化)と T5 転用の設計を主導(person / time-series)
- [[Lorenzo Stella]] — [[AWS AI Labs]]([[Amazon Web Services]])所属。Chronos([[@2024__arXiv__Chronos Learning the Language of Time Series]]、TMLR 2024)の equal contribution 筆頭著者(* 印)。[[Abdul Fatir Ansari]] と並ぶ共同筆頭著者(person / time-series)
- [[Yuyang Wang]] — [[AWS AI Labs]]([[Amazon Web Services]])所属。Chronos・[[Chronos-2]] の共同シニア著者(‡ 印)(person / time-series)
- [[Oleksandr Shchur]] — [[AWS AI Labs]]([[Amazon Web Services]])所属。Chronos・[[Chronos-2]] の共著者(person / time-series)
- [[Syama Sundar Rangapuram]] — [[Amazon Web Services]] 所属。[[Chronos-2]] 共著者。時系列確率予測・グローバル深層学習モデルの研究者。DeepAR 共同開発者(person / time-series)
- [[Danielle C. Maddix]] — [[Amazon Web Services]] 所属。[[Chronos-2]] 共著者。時系列基盤モデル・合成多変量データ生成(Multivariatizer)の設計に貢献(person / time-series)
- [[George Karypis]] — [[Amazon Web Services]] / University of Minnesota 所属。[[Chronos-2]] 共著者(シニア)。グラフ分割(METIS)・推薦システムでも著名(person / time-series / ml)
- [[Michael Bohlke-Schneider]] — [[Amazon Web Services]] 所属。[[Chronos-2]] 共同シニア著者(‡)。Amazon における時系列基盤モデル研究を Yuyang Wang と共同主導(person / time-series)
- [[Zhen Liu]] — [[South China University of Technology]] / [[Institute for Infocomm Research]] 所属。TSFM サーベイ [[@2026__techRxiv__From Pre-training to Post-training - A Survey on Time Series Foundation Models]] の co-first author(person / time-series)
- [[Qianli Ma]] — [[South China University of Technology]] 教授。TSFM サーベイ Liu+ 2026 の co-corresponding author。先行サーベイ「A survey on time-series pre-trained models」(TKDE 2024)も主導(person / time-series)
- [[Min Wu]] — [[Institute for Infocomm Research]](A*STAR Singapore)所属。TSFM サーベイ Liu+ 2026 の co-corresponding author(person / time-series)
- [[Eden Belouadah]] — [[Datadog]] AI Research 所属。Toto 2.0(arXiv:2605.20119)共著者(person / time-series)
- [[Marc Cenac]] — [[Datadog]] AI Research 所属。Toto 2.0 共著者(person / time-series)
- [[Xunyi Zhao]] — [[Datadog]] AI Research 所属。Toto 2.0 共著者(person / time-series)
- [[Viktoriya Zhukova]] — [[Datadog]] AI Research 所属。Toto 2.0 共著者(person / time-series)
- [[Othmane Abou-Amal]] — [[Datadog]] AI Research 所属。Toto 2.0 共著者(person / time-series)
- [[Youcef Remil]] — [[University of Lyon]] / [[INSA Lyon]] / [[CNRS]] UMR 5205 / [[Infologic]] 所属。[[@2024__arXiv__AIOps Solutions for Incident Management]] の筆頭著者(person / aiops)
- [[Anes Bendimerad]] — [[Infologic]] 所属。Remil+ 2024 AIOps サーベイ共著者(person / aiops)
- [[Romain Mathonat]] — [[Infologic]] 所属。Remil+ 2024 AIOps サーベイ共著者(person / aiops)
- [[Mehdi Kaytoue]] — [[University of Lyon]] / [[INSA Lyon]] / [[CNRS]] UMR 5205 / [[Infologic]] 所属。Remil+ 2024 AIOps サーベイ senior author。pattern mining 系の方法論主導者と推測(person / aiops / pattern-mining)
- [[Anson Bastos]] — [[Microsoft]] 所属。[[@2026__FSE__Attention Enhanced Entity Recommendation for Intelligent Monitoring in Cloud Systems]] 共著者。グラフ表現学習のスペクトル領域での表現力に関する先行研究(TMLR 2022)を持つ(person / aiops / graph-neural-network)
- [[Flora Salim]] — [[University of New South Wales]] の研究者。[[@2025__WWW Companion__RCAEval - A Benchmark for Root Cause Analysis of Microservice Systems with Telemetry Data]] 共著(person / aiops / microservices)
- [[Xiuzhen Zhang]] — [[RMIT University]] の研究者。[[@2025__WWW Companion__RCAEval - A Benchmark for Root Cause Analysis of Microservice Systems with Telemetry Data]] 共著(person / aiops / microservices)
- [[Yitao Yang]] — [[The Chinese University of Hong Kong]] の研究者。[[@2026__FSE__TSGuard - Automated User-Centric Incident Diagnosis for AI Workloads in the Cloud]] 筆頭著者(Microsoft Research インターン)(person / aiops / llm-agent)
- [[Yifan Xiong]] — [[Microsoft Research]] Vancouver の研究者。TSGuard 共著(person / aiops)
- [[Baochun Li]] — [[University of Toronto]] ECE の研究者。TSGuard 共著(person / aiops)
- [[Peng Cheng]] — [[Microsoft Research]] Redmond の研究者。TSGuard 共著(person / aiops)
- [[Walter Heimerdinger]] — [[Honeywell]] 所属。[[Charles Weinstock]] と共に SEI フォールトトレランス概念フレームワーク(CMU/SEI-92-TR-033、1992)を執筆(person / fault-tolerance)
- [[Charles Weinstock]] — [[Software Engineering Institute]]/[[Carnegie Mellon University]] の研究者。フォールトトレランス概念フレームワーク(CMU/SEI-92-TR-033、1992)の共著者(person / fault-tolerance)
- [[Felix Salfner]] — [[Humboldt University of Berlin]] の研究者。[[@2010__ACM CSUR__A Survey of Online Failure Prediction Methods]] 筆頭著者。hidden semi-Markov model (HSMM) によるオンライン障害予測の中心人物で、博士論文 "Event-based Failure Prediction" (2008) を公刊(person / dependability / failure-prediction)
- [[Maren Lenk]] — [[Humboldt University of Berlin]] 所属(2010 年時点)。[[@2010__ACM CSUR__A Survey of Online Failure Prediction Methods]] 2nd author(person / dependability)
- [[Miroslaw Malek]] — [[Humboldt University of Berlin]] 教授。ディペンダビリティ・プロアクティブ障害管理研究の指導的研究者。[[@2010__ACM CSUR__A Survey of Online Failure Prediction Methods]] senior author(person / dependability / failure-prediction)
- [[Patrick D. T. O'Connor]] — [[@2012__Wiley__Practical Reliability Engineering|Practical Reliability Engineering]] 第 5 版の著者。信頼性工学の定義、寿命データ解析、信頼性予測、DfR、FRACAS、信頼性管理を実務教科書として体系化(person / reliability)
- [[Andre Kleyner]] — [[@2012__Wiley__Practical Reliability Engineering|Practical Reliability Engineering]] 第 5 版の共著者。第 5 版でモンテカルロ、加速試験データ解析、保証データ解析、信頼性実証などを拡張(person / reliability)
- [[Claus Pahl]] — [[Free University of Bozen-Bolzano]] コンピュータサイエンス学部准教授。サービス・クラウドコンピューティングのソフトウェア工学。クラウドコンテナ技術 SMS(IEEE TCC 2019)対応著者(person / cloud / containers)
- [[Pooyan Jamshidi]] — [[Carnegie Mellon University]] ポスドク研究者(2017 時点)。高設定可能システム・機械学習・データ集約型計算。クラウドコンテナ技術 SMS 共著(person / cloud / ML)
- [[Paolo Notaro]] — [[TU Munich]] / [[Huawei Munich Research Center]] の研究者。AIOps for Failure Management サーベイ(ACM TIST 2021)筆頭著者(person / aiops / survey)
- [[Jorge Cardoso]] — [[University of Coimbra]] / [[Huawei Munich Research Center]] の研究者。AIOps for Failure Management サーベイ共著、Bento+ 2021(J Grid Computing)の Huawei Cloud OpenStack 本番トレース自動分析にも参画(person / aiops / tracing / survey)
- [[Andre Bento]] — [[University of Coimbra]] / CISUC の研究者。Bento+ 2021(J Grid Computing)の corresponding author。[[OpenTracing Processor]] (OTP) のリードオーサ(person / observability / tracing)
- [[Jaime Correia]] — [[University of Coimbra]] / CISUC の研究者。Bento+ 2021(J Grid Computing)と Pina+ 2018(IEEE NCA, Nonintrusive monitoring of microservice-based systems)の共著者(person / observability / microservices)
- [[Ricardo Filipe]] — [[University of Coimbra]] / CISUC の研究者。Bento+ 2021(J Grid Computing)と Pina+ 2018(IEEE NCA)の共著者(person / observability / microservices)
- [[Filipe Araujo]] — [[University of Coimbra]] / CISUC の研究者。Bento+ 2021(J Grid Computing)と Pina+ 2018(IEEE NCA)の共著者(person / observability / microservices)
- [[Michael Gerndt]] — [[TU Munich]] の Chair of Computer Architecture and Parallel Systems 教授。AIOps for Failure Management サーベイ共著(person / hpc / aiops)
- [[Satoru Kobayashi]] — [[University of Tokyo]] 大学院生(2018 年)。[[LogCausalAnalysis]] の主著者。PC アルゴリズム + G-square によるネットワーク syslog 因果マイニング(TNSM 2018)(person / network-management / causal-inference)
- [[Kazuki Otomo]] — [[University of Tokyo]] 大学院生(2018 年)。ネットワーク時系列知識抽出(TNSM 2018)(person / network-management)
- [[Kensuke Fukuda]] — [[National Institute of Informatics]] / SOKENDAI 准教授(2018 年)。インターネットトラフィック解析・異常検知(TNSM 2018)(person / network-management / anomaly-detection)
- [[Hiroshi Esaki]] — [[University of Tokyo]] 教授。JPNIC 副会長・WIDE Project 理事(2018 年)(person / network-management)
- [[University of Tokyo]] — 日本の国立総合大学。[[Satoru Kobayashi]]・[[Kazuki Otomo]]・[[Hiroshi Esaki]] の所属(organization / university)
- [[National Institute of Informatics]] — 日本の情報学国立研究機関(NII)。[[Kensuke Fukuda]] の所属(organization / research-institute)
- [[SINET4]] — 日本全国研究教育ネットワーク(800 以上の学術機関)。TNSM 2018 の評価データセット(456 日・35M 件 syslog)(dataset / network)
- [[LogCausalAnalysis]] — [[Satoru Kobayashi]] 公開の PC + G-square ネットワーク syslog 因果解析 OSS (repository / log-analysis)
- [[Xiang Rao]] — [[National University of Defense Technology]] 所属研究者。ノイズログフィルタリング手法 SBF(Haar ウェーブレット + DTW 類似度)の提案者。[[Alibaba Cloud]] との共同研究(person / log-analysis / distributed)
- [[Huaimin Wang]] — [[National University of Defense Technology]] 上席研究者。大規模分散システムの信頼性・障害診断分野(person / distributed / reliability)
- [[Francisco Neves]] — [[HASLab]]-INESC TEC / [[University of Minho]] 研究者。eBPF ブラックボックストラフィック監視・コンテナ配置最適化(SAC 2020)の筆頭著者(person / distributed / ebpf)
- [[Ricardo Vilaça]] — [[University of Minho]] 研究者。SAC 2020 コンテナ配置最適化論文の共著者(person / distributed)
- [[José Pereira]] — [[University of Minho]] 研究者。SAC 2020 コンテナ配置最適化論文の共著者(person / distributed)
- [[Tuomas Pelkonen]] — [[Facebook]] 所属エンジニア。[[Gorilla]] インメモリ TSDB(VLDB 2015)の筆頭著者(person / database / time-series)
- [[Marcus Müller]] — [[TU Munich]] 所属。B-Trees Are Back の筆頭著者(person / database)
- [[Lawrence Benson]] — [[TU Munich]] 所属。B-Trees Are Back の共著者(person / database)
- [[Viktor Leis]] — [[TU Munich]] 所属。B-Trees Are Back の共著者で、[[vmcache]] など DBMS 内部構造研究に関与(person / database)
- [[Lu Dai]] — [[Hong Kong University of Science and Technology]] 所属。LLM 向け情報検索・RAG ノイズ除去視点論文の筆頭著者
- [[Liang Sun]] — [[Hong Kong University of Science and Technology, Guangzhou]] 所属。LLM 向け情報検索論文の共著者
- [[Fanpu Cao]] — [[Hong Kong University of Science and Technology, Guangzhou]] 所属。LLM 向け情報検索論文の共著者
- [[Ziyang Rao]] — [[Hong Kong University of Science and Technology, Guangzhou]] 所属。LLM 向け情報検索論文の共著者
- [[Cehao Yang]] — [[Hong Kong University of Science and Technology, Guangzhou]] 所属。LLM 向け情報検索論文の共著者
- [[Hao Liu]] — [[Hong Kong University of Science and Technology, Guangzhou]] / [[Hong Kong University of Science and Technology]] 所属。LLM 向け情報検索論文の共著者
- [[Hui Xiong]] — [[Hong Kong University of Science and Technology, Guangzhou]] / [[Hong Kong University of Science and Technology]] 所属。LLM 向け情報検索論文の共著者
- [[Hengrui Wang]] — [[Tsinghua University]] 所属。[[EcoTune]] 論文の第一著者
- [[Jiansheng Qiu]] — [[Tsinghua University]] 所属。[[EcoTune]] 論文の共著者
- [[Fangzhou Yuan]] — [[Tsinghua University]] 所属。[[EcoTune]] 論文の共著者
- [[Huanchen Zhang]] — [[Tsinghua University]] 所属。[[EcoTune]] 論文の責任著者([[Shanghai Qi Zhi Institute]] 兼務)
- [[Dan Hendrycks]] — [[Center for AI Safety]] 創設者・エグゼクティブディレクター。MMLU 設計者・HLE 上級著者([[@2025__arXiv__Humanity's Last Exam]])
- [[Long Phan]] — [[Center for AI Safety]] 所属。HLE 共同第一著者([[@2025__arXiv__Humanity's Last Exam]])
- [[Suman Karumuri]] — [[Slack Technologies]] 所属エンジニア。[[オブザーバビリティデータモデル|ODMS]] ビジョン論文の筆頭著者(SIGMOD Record 2021)
- [[Franco Solleza]] — [[Brown University]] 所属研究者。ODMS ビジョン論文の共著者
- [[Stan Zdonik]] — [[Brown University]] 教授。ストリーム処理・時系列データ管理の専門家
- [[Nesime Tatbul]] — Intel Labs・MIT 所属研究者。時系列データ管理・異常検知の専門家
- [[Tom Henighan]] — [[OpenAI]] 所属。スケーリング則多モダリティ拡張論文(arXiv:2010.14701)の均等貢献筆頭著者。画像・動画実験を担当
- [[Pooja Srinivas]] — [[Microsoft]] 所属。インテリジェント監視フレームワーク(ICSE-SEIP 2024)の第一著者
- [[Fiza Husain]] — [[Microsoft]] 所属。インテリジェント監視フレームワーク(ICSE-SEIP 2024)の共著者
- [[Ayush Choure]] — [[Microsoft]] 所属。インテリジェント監視フレームワーク(ICSE-SEIP 2024)の共著者
- [[Avi Nayak]] — [[Microsoft]] Senior Program Manager。Intelligent Monitoring Framework プロジェクト参加(MSR Blog 2024)
- [[Piyali Jana]] — [[Microsoft]] Principal Software Engineer Manager。Intelligent Monitoring Framework プロジェクト参加(MSR Blog 2024)
- [[Daemyung Kang]] — [[Lablup Inc]] の 504 GPU 本番訓練クラスタ運用分析論文の筆頭著者
- [[Wei Bai]] — [[Microsoft]] 所属。[[Azure Storage]] のリージョン内 RDMA 展開経験論文の筆頭著者
- [[Dmitry Arkhangelskiy]] — [[Amazon Web Services]] 所属。[[Aurora Limitless Database]] 論文の筆頭著者
- [[Jacopo Soldani]] — [[University of Pisa]] 所属。マイクロサービス異常検知・RCA 統合サーベイ(ACM CSUR 2021)の筆頭著者
- [[Antonio Brogi]] — [[University of Pisa]] 所属。マイクロサービス異常検知・RCA 統合サーベイ(ACM CSUR 2021)の共著者
- [[Luís M. Barata]] — [[Instituto de Telecomunicações]] / [[Universidade da Beira Interior]] / [[Instituto Politécnico de Castelo Branco]] 所属。マイクロサービス異常検知・根本原因特定サーベイの筆頭著者
- [[Sérgio Sequeira]] — [[Universidade da Beira Interior]] 所属。マイクロサービス異常検知・根本原因特定サーベイの共著者
- [[Eurico Lopes]] — [[Instituto Politécnico de Castelo Branco]] 所属。マイクロサービス異常検知・根本原因特定サーベイの共著者
- [[Pedro R. M. Inácio]] — [[Instituto de Telecomunicações]] / [[Universidade da Beira Interior]] 所属。マイクロサービス異常検知・根本原因特定サーベイの共著者
- [[Mário M. Freire]] — [[Instituto de Telecomunicações]] / [[Universidade da Beira Interior]] / [[NOVA LINCS]] 所属。マイクロサービス異常検知・根本原因特定サーベイの共著者
- [[Zefan Wang]] — [[Tsinghua University]] 所属。[[RCAgent]] 論文の第一著者
- [[Zichuan Liu]] — [[Nanjing University]] 所属。[[RCAgent]] 論文の共同第一著者
- [[Yingying Zhang]] — [[Alibaba Group]] 所属。[[RCAgent]] 論文の責任著者
- [[Aoxiao Zhong]] — [[Harvard University]] 所属。[[RCAgent]] 論文の共著者
- [[Jihong Wang]] — [[Xi’an Jiaotong University]] 所属。[[RCAgent]] 論文の共同第一著者
- [[Fengbin Yin]] — [[Alibaba Group]] 所属。[[RCAgent]] 論文の共著者
- [[Lunting Fan]] — [[Alibaba Group]] 所属。[[RCAgent]] 論文の共著者
- [[Lingfei Wu]] — [[Anytime AI]] 所属。[[RCAgent]] 論文の共著者
- [[Qingsong Wen]] — [[Squirrel Ai Learning]] 所属。[[RCAgent]] 論文の責任著者
- [[Vaishali Vinay]] — [[Microsoft]] Security Research 所属。LLM アプリケーション失敗モードのシステムレベルタクソノミーを提示した IEEE CAI 2026 論文の著者
- [[Liz Fong-Jones]] — [[Honeycomb.io|Honeycomb]] Principal Developer Advocate。オブザーバビリティ・SRE コミュニティの著名人。[[CNCF TAG Observability Whitepaper]](2023)の主要貢献者。
- [[赤穂昭太郎]] — [[産業技術総合研究所]] 人間情報インタラクション研究部門 上級主任研究員。[[統計的機械学習]]・[[ベイズ最適化]] の研究者
- [[産業技術総合研究所]] — 国立研究開発法人(AIST)。日本最大規模の公的研究機関の一つ
- [[Kuaishou Technology]] — 中国短動画プラットフォーム企業。[[Bian Que]] エージェント型 O&M フレームワーク開発元
- [[Bian Que]] — [[Kuaishou Technology]] 開発のエージェント型 O&M フレームワーク。統一運用パラダイム・[[Flexible Skill Arrangement]]・統一自己進化メカニズムで構成
- [[Bochao Liu]] — [[Kuaishou Technology]] 所属。[[Bian Que]](arXiv:2604.26805)の均等貢献第一著者
- [[Ben Chen]] — [[Kuaishou Technology]] 所属。[[Bian Que]](arXiv:2604.26805)の責任著者
- [[Nirmesh Malviya]] — [[MIT CSAIL]] 所属の研究者(2014 年当時)。コマンドロギングと ARIES 生理ロギングの詳細比較を行った ICDE 2014 論文の筆頭著者([[@2014__ICDE__Rethinking Main Memory OLTP Recovery]])(person / database / recovery)
- [[Ariel Weisberg]] — [[VoltDB]] Inc. エンジニア・研究者(2014 年当時)。VoltDB にコマンドロギングと生理ロギングの両実装を担当した ICDE 2014 論文の共著者([[@2014__ICDE__Rethinking Main Memory OLTP Recovery]])(person / database / recovery)
- [[MIT CSAIL]] — MIT のコンピュータ科学・人工知能研究所。[[Michael Stonebraker]]・[[Samuel Madden]]・[[Nirmesh Malviya]] が所属し、[[H-Store]]・[[VoltDB]] の設計研究を主導(organization / university / database)
- [[VoltDB]] — [[H-Store]] の設計を商用化したオープンソースのメインメモリ分散 OLTP データベース。コマンドロギング・ノンブロッキングチェックポイント・パーティション単位逐次実行を採用(product / database / OLTP)
- [[Michael Stonebraker]] — Turing Award 受賞者。データベース分野の先駆的研究者。「One Size Fits All」(ICDE 2005)で専用データベースシステムの必要性を主張し、H-Store(VLDB 2007)で OLTP 向けメインメモリ DB を実証
- [[Ugur Cetintemel]] — [[Brown University]] 准教授。「One Size Fits All」(ICDE 2005)の共著者
- [[Jeffrey Dean]] — [[Google]] フェロー。[[Bigtable]](OSDI 2006)の筆頭著者。MapReduce・TensorFlow 等の共同設計者
- [[Sanjay Ghemawat]] — [[Google]] フェロー。Bigtable(OSDI 2006)の共著者。GFS・MapReduce の共同設計者
- [[Fabio Baltieri]] — [[Google]] 所属。[[Bigtable]] の 20 年運用経験論文([[@2026__SIGMOD Companion__Twenty Years of Bigtable]])の筆頭著者
- [[Werner Vogels]] — [[Amazon]] CTO。[[Dynamo]](SOSP 2007)の最終著者・対外的な技術スポークスパーソン
- [[Giuseppe DeCandia]] — Amazon の分散システムエンジニア。[[Dynamo]](SOSP 2007)の筆頭著者
- [[Avinash Lakshman]] — [[Dynamo]] の共著者(Amazon)→ [[Apache Cassandra]] の筆頭著者(Facebook)。2 つの分散ストレージを接続する人物
- [[Prashant Malik]] — [[Facebook]] エンジニア。Cassandra(LADIS 2009 / SIGOPS OSR 2010)の共著者
- [[Samuel Madden]] — [[MIT]] 教授。H-Store(VLDB 2007)の共著者。データベースシステム研究
- [[Daniel J. Abadi]] — H-Store(VLDB 2007)の共著者。カラムストア C-Store の共同設計者
- [[Stavros Harizopoulos]] — H-Store(VLDB 2007)の共著者。データベースアーキテクチャ研究
- [[Pat Helland]] — H-Store(VLDB 2007)の共著者。分散トランザクションシステムの著名研究者(元 Tandem/Microsoft/Amazon)
- [[Wei-Lin Chiang]] — [[University of California, Berkeley]] 博士課程。[[LMSYS]] 共同設立者。[[Chatbot Arena]] 第一著者(同等貢献、arXiv 2024)
- [[Lianmin Zheng]] — [[University of California, Berkeley]] 博士課程。[[LMSYS]] 共同設立者。[[Chatbot Arena]] 第一著者(同等貢献、arXiv 2024)。SGLang・Vicuna の主要著者
- [[Katja Gilly]] — [[Miguel Hernández University]] 所属研究者。ウェブロードバランシングサーベイ([[@2011__World Wide Web__An up-to-date survey in web load balancing]])の筆頭著者(person / distributed / web-systems)
- [[Carlos Juiz]] — [[University of Balearic Islands]] 所属研究者。ウェブロードバランシングサーベイの共著者(person / distributed / web-systems)
- [[Ramon Puigjaner]] — [[University of Balearic Islands]] 所属研究者。ウェブロードバランシングサーベイの共著者(person / distributed / web-systems)
## Organization
- [[Topotal]] — [[Ryota Yoshikawa]] が CTO を務める企業。SRE as a Service と、AI も活用したインシデントマネジメント SaaS [[Waroom]] を扱う(organization / sre / incident-management)
- [[Meta FAIR]] — Meta AI Research(旧 Facebook AI Research)。[[Zeyuan Allen-Zhu]]・[[Yann LeCun]] が所属し、[[Physics of Language Models]] シリーズや次トークン予測の限界に関する研究を発信([[joisino-言語モデルの物理学-2025]]、[[joisino-LLMと言葉の感じ方-2026]])(organization / industrial-lab / machine-learning)
- [[Anthropic]] — Claude を提供する AI スタートアップ。[[joisino-否定文理解-2024]] では[[文脈付き検索]](Contextual Retrieval、5.0%→2.9% 検索ミス削減)、[[joisino-人間を騙すAI-2025]] では[[報酬ハッキング]]・[[スコファンシ]]の RLHF サーベイ源として参照される(organization / ai-startup / safety-research)
- [[Sun Microsystems]] — カリフォルニア州サンタクララのコンピュータメーカー(2010 年 Oracle に買収)。Solaris OS の開発元で [[DTrace]] の発祥地。[[Bryan Cantrill]]・[[Michael Shapiro]]・[[Adam Leventhal]] が Solaris カーネルに DTrace を統合した(organization / systems)
- [[LBNL]] — Lawrence Berkeley National Laboratory。米国エネルギー省傘下の国立研究所(カリフォルニア州バークレー)。[[Steven McCanne]]・[[Van Jacobson]] が [[BPF]] を開発した研究拠点(organization / research-lab / networking)
- [[Hasso Plattner Institute]] — ドイツ・ポツダム大学付属の情報技術研究機関(HPI)。Felix Naumann のデータ品質グループと [[Thorsten Papenbrock]]・[[Phillip Wenig]]・[[Sebastian Schmidl]] が [[TimeEval]]・[[GutenTAG]] を開発(organization / university / research)
- [[NYU Abu Dhabi]] — アラブ首長国連邦アブダビのニューヨーク大学キャンパス。[[Djellel Difallah]] が所属しデータベースベンチマーク・クラウドソーシング研究を行う(organization / university)
- [[eXascaleInfolab]] — [[University of Fribourg]](スイス)の時系列・エクサスケールデータ管理研究グループ。[[Philippe Cudré-Mauroux]] 主宰。ORBITS・Mind the Gap・TSM-Bench 等の時系列研究を主導(organization / research-group / time-series)
- [[South China University of Technology]] — 中国・広州市の国立研究大学(華南理工大学、SCUT)。School of Computer Science and Engineering は時系列基盤モデル研究の中核グループのひとつで、[[Qianli Ma]]・[[Zhen Liu]] が事前学習・事後学習サーベイを主導(organization / university / time-series)
- [[Institute for Infocomm Research]] — シンガポール A*STAR 傘下の公立研究機関(I2R)。時系列・グラフ機械学習で [[Min Wu]]・[[Zhen Liu]] が活動し、TSFM サーベイ Liu+ 2026 を共同主導(organization / research-institute / time-series)
- [[University of Lyon]] — フランス・リヨンの大学連合体(Université de Lyon)。[[INSA Lyon]] と [[CNRS]] UMR 5205(LIRIS)と協働(organization / university)
- [[INSA Lyon]] — フランス・ヴィルールバンヌの工学系大学(Institut National des Sciences Appliquées de Lyon)。[[University of Lyon]] 連合・[[CNRS]] UMR 5205 と協働。[[Youcef Remil]]・[[Mehdi Kaytoue]] の所属(organization / university)
- [[CNRS]] — フランス公的研究機関(Centre national de la recherche scientifique)。UMR 5205(LIRIS)は [[University of Lyon]] / [[INSA Lyon]] との共同研究室で、データマイニング・パターン発見の拠点(organization / research-institute)
- [[Infologic]] — フランス・ブール=レ=ヴァランス(FR-26500)の企業。AIOps 産学連携の本拠で [[University of Lyon]] / [[INSA Lyon]] / [[CNRS]] UMR 5205 と協働。Remil+ 2024 AIOps サーベイで Infologic SQL Queries / Alerts / ASH を public dataset として提供(organization / industry / aiops)
- [[University of New South Wales]] — オーストラリア・シドニーの公立研究大学(UNSW)。[[Flora Salim]] 所属。[[RCAEval]] 共著機関(organization / university)
- [[Microsoft Research]] — [[Microsoft]] の研究部門(MSR)。Redmond と Vancouver の研究員 [[Peng Cheng]]・[[Yifan Xiong]] が CUHK と共同で [[TSGuard]] を構築。Azure 本番運用と連携した AIOps/インシデント研究を推進(organization / industry-lab / aiops)
- [[University of Toronto]] — カナダ・トロントの公立研究大学。ECE の [[Baochun Li]] が [[TSGuard]] 共著(organization / university)
- [[Honeywell]] — 米国の産業コングロマリット。航空宇宙・制御システム・セキュリティに強みを持つ。[[Walter Heimerdinger]] の所属組織(organization / industry)
- [[Software Engineering Institute]] — CMU 内の米国防省資金による連邦政府資金研究開発センター(FFRDC)。ソフトウェアエンジニアリング・セキュリティ研究機関。[[Charles Weinstock]] の所属(organization / research-institute)
- [[Humboldt University of Berlin]] — ベルリンの研究大学(Humboldt-Universität zu Berlin)。コンピュータサイエンス学部に [[Miroslaw Malek]] らが率いる dependable computing グループを擁し、HSMM ベースのオンライン障害予測研究とサーベイ([[@2010__ACM CSUR__A Survey of Online Failure Prediction Methods]])の発信源(organization / university / dependability)
- [[Wiley]] — John Wiley & Sons。[[@2012__Wiley__Practical Reliability Engineering|Practical Reliability Engineering]] 第 5 版の出版元(organization / publisher)
- [[Free University of Bozen-Bolzano]] — イタリア・ボルツァーノ所在の大学。[[Claus Pahl]] 所属。Edge Cloud Orchestration プロジェクトを擁する(organization / university)
- [[University of Pisa]] — イタリア・ピサ所在の大学。[[Antonio Brogi]]・[[Jacopo Soldani]] 所属の SOCC(Service-Oriented and Cloud Computing)研究グループを擁する。PRA 2016 64 Through the fog プロジェクト(organization / university)
- [[University of Coimbra]] — ポルトガル中部 Coimbra に拠点を置く研究大学(Universidade de Coimbra)。Department of Informatics Engineering / CISUC を擁し [[Jorge Cardoso]] が所属(organization / university)
- [[Huawei Munich Research Center]] — [[Huawei Technologies]] のミュンヘン研究所(Riessstr. 25, 80992 München)。[[Paolo Notaro]]・[[Jorge Cardoso]] の研究拠点(organization / industry-lab)
- [[National University of Defense Technology]] — 中国・湖南省長沙市に所在する国防省立大学(NUDT / 国防科技大学)。国家平行分散処理重点実験室を擁し、分散システム・ログ解析・障害診断分野を研究(organization / university)
- [[HASLab]] — ポルトガルの INESC TEC 傘下、[[University of Minho]] 附属の分散システム研究ラボ(organization / distributed)
- [[University of Minho]] — ポルトガル Braga 所在の大学。[[Francisco Neves]]・[[Ricardo Vilaça]]・[[José Pereira]] が所属(organization / university)
- [[Shanghai Qi Zhi Institute]] — [[EcoTune]] 論文で [[Huanchen Zhang]] の兼務所属として記載される研究機関
- [[Center for AI Safety]] — AI 安全性研究の非営利機関。[[Dan Hendrycks]] が創設。HLE ベンチマーク([[@2025__arXiv__Humanity's Last Exam]])を [[Scale AI]] と共同開発
- [[Scale AI]] — AI データラベリング・評価プラットフォーム企業。HLE 論文の主要所属機関([[@2025__arXiv__Humanity's Last Exam]])
- [[CNCF]] — Cloud Native Computing Foundation。Linux Foundation 傘下の非営利組織。Kubernetes・Prometheus・[[OpenTelemetry]] など主要クラウドネイティブ OSS をホスト。[[TAG Observability]] など複数の Technical Advisory Group を擁する。
- [[TAG Observability]] — CNCF の Technical Advisory Group for Observability。35+ 名の貢献者による[[オブザーバビリティ]]ホワイトペーパー(v1.0、2023 年 10 月)の策定組織。Source: [[@2023__CNCF TAG Observability__Observability Whitepaper]]。
- [[Lablup Inc]] — [[Backend.AI]] と [[Sokovan]] を開発する AI インフラ企業。504 GPU 本番訓練クラスタ運用分析を報告
- [[Instituto de Telecomunicações]] — [[Luís M. Barata]]・[[Pedro R. M. Inácio]]・[[Mário M. Freire]] の所属として登場するポルトガルの研究機関
- [[Universidade da Beira Interior]] — [[Luís M. Barata]]・[[Sérgio Sequeira]]・[[Pedro R. M. Inácio]]・[[Mário M. Freire]] の所属として登場するポルトガルの大学
- [[Instituto Politécnico de Castelo Branco]] — [[Luís M. Barata]]・[[Eurico Lopes]] の所属として登場するポルトガルの高等教育機関
- [[NOVA LINCS]] — [[Mário M. Freire]] の所属として登場する NOVA Laboratory for Computer Science and Informatics
- [[Cluster Computing]] — Springer Nature 系の論文誌。[[Luís M. Barata]] らのマイクロサービス異常検知・根本原因特定サーベイを掲載
- [[Xi’an Jiaotong University]] — [[RCAgent]] 論文の著者所属
- [[Teradata]] — 1978 年からシェアードナッシング・コモディティ CPU・SQL の組み合わせによる並列データベースを商業化した企業。IFP/AMP プロセッサ分割・Y-net デュアルインターコネクト・ハッシュパーティショニングを採用し 1992 年時点で 1,000+ プロセッサを出荷。(organization / database / parallel)
- [[Anytime AI]] — [[RCAgent]] 論文の著者所属
- [[Squirrel Ai Learning]] — [[RCAgent]] 論文の著者所属
- [[MIT]] — マサチューセッツ工科大学。H-Store の主要研究拠点(Samuel Madden 所属)
- [[Brown University]] — ブラウン大学。Michael Stonebraker の所属。H-Store プロジェクトの拠点
- [[Amazon]] — Amazon.com。高可用キーバリューストア [[Dynamo]](SOSP 2007)を開発
- [[Facebook]] — ソーシャルメディア企業(現 Meta)。[[Apache Cassandra]] を開発し Inbox Search に投入
- [[LMSYS]] — Large Model Systems Organization。[[University of California, Berkeley]] 主導の LLM 研究グループ。[[Chatbot Arena]]・Vicuna・SGLang・LMSYS-Chat-1M を開発・公開
- [[Miguel Hernández University]] — スペイン・アリカンテ州エルチェ所在の公立大学(UMH)。[[Katja Gilly]] の所属機関(organization / university)
- [[University of Balearic Islands]] — スペイン・バレアレス諸島州パルマ所在の公立大学(UIB)。[[Carlos Juiz]]・[[Ramon Puigjaner]] の所属機関(organization / university)
## Product / System
- [[Waroom]] — [[Topotal]] が開発する、AI も活用したインシデントマネジメント SaaS(product / incident-management / sre)
- [[TimeEval]] — [[Hasso Plattner Institute]] の [[Phillip Wenig]]・[[Sebastian Schmidl]]・Thorsten Papenbrock が開発した時系列異常検知の分散評価フレームワーク。71 手法を Docker コンテナ化し 976 データセットで自動評価。[[@2022__PVLDB__Anomaly Detection in Time Series - A Comprehensive Evaluation]] の評価基盤。OSS: https://github.com/HPI-Information-Systems/TimeEval(product / anomaly-detection / time-series / benchmark)
- [[GutenTAG]] — [[Hasso Plattner Institute]] が開発した合成時系列データ生成ツール。異常型・ノイズ・周期性を制御可能なラベル付き時系列を生成し、時系列異常検知ベンチマークのラベル品質問題を解決する。[[@2022__PVLDB__Anomaly Detection in Time Series - A Comprehensive Evaluation]] で活用。OSS: https://github.com/HPI-Information-Systems/gutentag(product / anomaly-detection / time-series / data-generation)
- [[TSGuard]] — AI ワークロード向けの user-centric 多エージェントインシデント診断システム。CUHK・MSR・UoT が Microsoft Azure 本番データで構築。quick path(履歴 DB)→ slow path(タクソノミー DFS)→ deep path(探索)の 3 段パイプライン + 5 エージェント協調(product / aiops / llm-agent)
- [[RCACopilot]] — Microsoft 系の LLM ベースクラウドインシデント RCA システム(Chen+ EuroSys 2024)。GPT-4o + RAG で one-shot 推論。TSGuard 等のベースライン(product / aiops / rca)
- [[OpenTracing]] — 分散トレーシングのベンダー中立 API/抽象標準。trace ID・span ID・親 span ID で span tree を構成、annotation で key-value メタ情報を付与。OpenTracing Specification Council が管理し、2019 年に OpenCensus と統合する形で [[OpenTelemetry]] に合流(product / observability / tracing)
- [[Docker]] — Linux LXC のメカニズム(namespace・cgroup)を利用したコンテナ事実上の標準ソリューション。union mount でレイヤー化された読み書き可能ファイルシステムを構築(product / container)
- [[LXC]] — Linux Container プロジェクト(2007 年頃)。namespace・cgroup を提供しコンテナ研究の起点。Docker・systemd-nspawn 等の基盤(product / container)
- [[Gorilla]] — [[Facebook]] が開発・本番展開したインメモリ時系列データベース(TSDB)。デルタ・オブ・デルタ + XOR 圧縮で 12 倍圧縮、HBase 比クエリレイテンシ 73 倍削減。(product / time-series)
- [[vmcache]] — Virtual-Memory Assisted Buffer Management に基づく storage engine / buffer management 基盤。B-Trees Are Back の system integration 評価で使用
- [[EcoTune]] — LSM ツリーの平均クエリスループットを最適化する動的計画法ベースのコンパクション方針
- [[RocksDB]] — LSM ツリーベースのキーバリューストア。[[EcoTune]] の評価実装対象
- [[Backend.AI]] — [[Lablup Inc]] の AI/ML ワークロード管理プラットフォーム。セッション単位の訓練ライフサイクルと自動リトライを提供
- [[Sokovan]] — [[Backend.AI]] の GPU 中心スケジューリング層。NUMA-aware 配置と 60 ノード訓練のギャングスケジューリングを担う
- [[Azure Storage]] — [[Microsoft]] Azure のクラウドストレージサービス。計算とストレージを分離し、sU-RDMA/sK-RDMA によるフロントエンド/バックエンド通信を導入
- [[RDMA Estats]] — [[Azure Storage]] の RDMA 展開で用いられたホスト側診断テレメトリ。RDMA 操作レイテンシをホスト/NIC/ネットワークに分解
- [[Aurora Limitless Database]] — [[Amazon Web Services]] の分散 OLTP データベースシステム。Amazon Aurora PostgreSQL をルータ/シャード構成へ拡張し、PostgreSQL 互換性と強い整合性を維持した水平スケーリングを狙う
- [[RCAgent]] — クラウド RCA 向けツール拡張 LLM 自律エージェントフレームワーク。[[Alibaba Cloud]] の Apache Flink OoD ジョブ診断に統合
- [[Dynamo]] — [[Amazon]] の内部キーバリューストア。結果整合性・一貫性ハッシュ法・ベクタークロック・スロッピークォーラムを組み合わせた高可用設計(SOSP 2007)
- [[Bigtable]] — [[Google]] の分散ストレージ/非リレーショナルデータベース。OSDI 2006 の多次元疎マップ設計から、SIGMOD Companion 2026 時点で 10 EB・ピーク 70 億 QPS 規模へ成長
- [[Google File System]] — Google の分散ファイルシステム(GFS)。Bigtable の永続化層
- [[Chubby]] — Google の分散ロックサービス。Bigtable のタブレット割り当て・スキーマ管理に使用
- [[Apache Cassandra]] — Dynamo のパーティショニングと Bigtable のカラムファミリモデルを統合した分散ストレージ(Facebook 発、現 Apache プロジェクト)
- [[H-Store]] — メインメモリ・単一スレッド・シェアードナッシングの OLTP プロトタイプ。商用 RDBMS 比 82 倍のスループット(VLDB 2007)。後継 [[VoltDB]] の設計基盤(product / database / OLTP)
- [[Chatbot Arena]] — [[LMSYS]] 開発の LLM 評価オープンプラットフォーム。匿名ペアワイズ比較 + Bradley-Terry モデルによるクラウドソーシング型ランキング。2024 年 1 月時点 240K 票・90K ユーザー
## Person
- [[Qinghao Hu]] — LLM 開発ワークロード特徴づけ記事の著者。[[Acme]] 6 か月トレースから LLM 専用 GPU クラスタのジョブ分布・利用率・失敗影響を分析
- [[Tianwei Zhang]] — LLM 開発ワークロード特徴づけ記事の著者。[[Qinghao Hu]]・[[Peng Sun]] とともに [[Acme]] の 6 か月トレースを分析
- [[Apostolos Kokolis]] — Meta の ML 研究クラスタ信頼性論文(HPCA 2025)の共同筆頭著者
- [[Michael Kuchnik]] — Meta の ML 研究クラスタ信頼性論文(HPCA 2025)の共同筆頭著者
- [[Carole-Jean Wu]] — Meta の ML システム研究者。Revisiting Reliability 論文のシニア著者
- [[Myeongjae Jeon]] — Philly トレース論文(USENIX ATC 2019)の筆頭著者([[UNIST]] / [[Microsoft]])
- [[Shivaram Venkataraman]] — Philly トレース論文(USENIX ATC 2019)の共著者([[University of Wisconsin]] / [[Microsoft]])
- [[Amar Phanishayee]] — Philly トレース論文(USENIX ATC 2019)の共著者([[Microsoft]])
- [[Junjie Qian]] — Philly トレース論文(USENIX ATC 2019)の共著者([[Microsoft]])
- [[Wencong Xiao]] — Philly トレース論文(USENIX ATC 2019)の共著者([[Beihang University]] / [[Microsoft]])、後続の大規模訓練インフラ研究にも登場
- [[Fan Yang]] — Philly トレース論文(USENIX ATC 2019)の共著者([[Microsoft]])
- [[Adrián Pérez Diéguez]] — Qualcomm Technologies 所属。LLM 事前学習性能チューニング論文(PMBS25)の筆頭著者
- [[Guibin Zhang]] — NUS 研究者。Agentic RL サーベイ(TMLR 2026)の筆頭著者
- [[Heng Ji]] — UIUC 教授。自然言語処理の著名研究者。Agentic RL サーベイのシニア著者
- [[Mi Zhang]] — Ohio State University 准教授、AIoT-MLSys Lab 主宰。Efficient LLMs サーベイ(TMLR 2024)の責任著者
- [[Mosharaf Chowdhury]] — University of Michigan 准教授。ML システム・ネットワーク研究(Gavel/Tiresias/Perseus 等)
- [[Pieter Hijma]] — [[Vrije Universiteit Amsterdam]] 准教授。GPU プログラミング最適化の体系的文献調査([[@2023__CSUR__Optimization Techniques for GPU Programming]])の筆頭著者。28 技術・4 テーマ分類体系を確立
- [[Stijn Heldens]] — [[Netherlands eScience Center]] 研究者。GPU 最適化 CSUR 論文([[@2023__CSUR__Optimization Techniques for GPU Programming]])の共著者。litstudy 文献解析ライブラリの開発者
- [[Ben van Werkhoven]] — [[Netherlands eScience Center]] 研究者。[[Kernel Tuner]] 汎用 GPU カーネル auto-tuning フレームワークの開発者。GPU 最適化 CSUR 論文共著
- [[Henri E. Bal]] — [[Vrije Universiteit Amsterdam]] 教授。GPU 最適化 CSUR 論文のシニア著者。高性能コンピューティング・分散システム研究
- [[Alessio Sclocco]] — [[Netherlands eScience Center]] 研究者。GPU 最適化 CSUR 論文([[@2023__CSUR__Optimization Techniques for GPU Programming]])の共著者
## Organization
- [[Slack Technologies]] — クラウドベースビジネスコミュニケーションプラットフォーム(Salesforce 傘下)。MELT 観測データ管理の産業事例研究の主題。2020年時点で Metrics 12M samples/秒・Events 250TB/日・Logs 90TB/日・Traces 2TB/日を運用
- [[UNIST]] — Philly トレース論文(USENIX ATC 2019)の著者所属
- [[University of Wisconsin]] — Philly トレース論文(USENIX ATC 2019)の著者所属
- [[Qualcomm]] — LLM 事前学習性能チューニング論文(PMBS25)の著者所属組織
- [[The Ohio State University]] — 米国オハイオ州コロンバスの州立研究大学。AIoT-MLSys Lab が Efficient LLMs サーベイを主導
- [[Vrije Universiteit Amsterdam]] — オランダの公立研究大学(VU Amsterdam)。[[Pieter Hijma]]・[[Henri E. Bal]] の所属機関。GPU プログラミング最適化 CSUR 論文の主要研究拠点
- [[Netherlands eScience Center]] — オランダの eScience 研究機関(NLeSC)。[[Stijn Heldens]]・[[Ben van Werkhoven]]・[[Alessio Sclocco]] の所属機関。[[Kernel Tuner]] 開発元
- [[University of Copenhagen]] — デンマーク・コペンハーゲンの総合研究大学。[[Yongluan Zhou]] の所属機関。[[ByteSeries]](SoCC 2020)の著者所属
## Product / System
- [[Acme]] — [[Shanghai AI Laboratory]] の LLM 開発向け GPU データセンター。Seren/Kalos の 2 クラスタ、計 4,704 A100 GPU
- [[InternEvo]] — [[Shanghai AI Laboratory]] の LLM 事前学習フレームワーク。V2 が 123B LLM・2,048 GPU で V1 比約 16% 高速化
- [[Philly]] — Microsoft の DNN 訓練向けマルチテナント GPU クラスタ管理サービス。YARN Fair Scheduler を基盤にギャングスケジューリングと局所性考慮配置を行う
- [[Meta AI Research SuperCluster]] — Meta の A100 世代 ML 研究クラスタ群。RSC-1 は 16k GPU、RSC-2 は 8k GPU
- [[NIXL]] — NVIDIA が開発した LLM 推論向け高帯域・低レイテンシデータ転送ライブラリ。NB API/SB API/バックエンドプラグインの3層構造で UCX・GDS・OBJ・MoonCake・HF3FS の5バックエンドをサポート。NVIDIA Dynamo の構成要素
- [[UCX]] — Unified Communication X。高帯域・低レイテンシネットワーク向けオープンソース通信フレームワーク。UCT(トランスポート抽象化)と UCP(ビジネスロジック)の2層構造。NIXL のデフォルトバックエンド
- [[CASCA]] — TU Wien のカーボン認識 SLO 管理マイクロサービスプラットフォーム(EMMa + RLDS、MSA 原則準拠のプライバシー保護設計、arXiv:2602.12875, 2026)
- [[TAMO]] — Shandong University の [[Xiao Zhang]]・[[Dongxiao Yu]] らによるツール支援型 LLM マルチモーダル RCA フレームワーク(双分岐拡散 T1 + FFT+GAT T2 + Transformer T3 + GPT-4 エージェント、IEEE TSC 2025)
- [[Azure Copilot]] — Azure 向けに調整された GPT-4 ベース LLM(クラウド管理ビジョン論文の SDK/CLI/IaC エージェントのモデル)
- [[WorkArena]] — Web エージェントの knowledge work ベンチマーク/実装基盤(クラウド管理ビジョン論文の ClickOps エージェントの土台、GPT-4o)
- [[SAKURAONE]] — SAKURA Internet の 800 GPU(100×H100)オープン Ethernet AI–HPC クラスタ(SONiC+RoCEv2 の 800 GbE、TOP500 HPL 49 位・トップ 100 唯一のフルオープンなネットワーキングスタック)
- [[SONiC]] — オープンソースのネットワーク OS。SAI で switching ASIC を抽象化(RoCEv2 のロスレス Ethernet を提供、SAKURAONE 採用)
- [[MegaScale]] — ByteDance/PKU の 10,000 GPU 超 LLM 訓練本番システム(175B を 12,288 GPU・55.2% MFU、Megatron-LM 基盤)
- [[Minder]] — ByteDance の大規模分散訓練向け自動の障害マシン検出器(マシン単位の類似度 + 連続性 + メトリクスごとの LSTM-VAE、本番 1 年超稼働、precision 0.904)
- [[Megatron-LM]] — NVIDIA の SOTA な OSS LLM 訓練フレームワーク(3D parallelism、MegaScale の基盤兼ベースライン)
- [[Pulse]] — Nanjing University の LLM 訓練トラフィック中心の監視システム(BlueField-3 上のマイクロ秒 RDMA 計測 → オペレータのセグメンテーション → マシン単位の箇所特定、非侵入的、12 中 10)
- [[NCCL]] — NVIDIA の集団通信ライブラリ。LLM 分散訓練の標準 CCL(ring/tree collective・P2P・all-to-all、Pulse のフック対象)
- [[BlueField-3]] — NVIDIA のプログラマブルな SmartNIC/DPU(DPA マイクロプロセッサ + 1GB DDR)。Pulse NIC Agent のマイクロ秒 RDMA 計測プラットフォーム
- [[Aegis]] — Alibaba の AI 訓練クラウド向け障害診断(OP-level・CCL 改変、Pulse の SOTA ベースライン、NSDI 25)
- [[Holmes]] — メガスケール GPU クラスタの LLM 訓練の不規則性の箇所特定(OP-level・ストラグラー特化、Pulse の SOTA ベースライン、NSDI 25)
- [[GreyHound]] — hybrid-parallel 訓練の fail-slow 検出(非侵入的だが OP-level、Pulse の SOTA ベースライン、ATC 25)
- [[SREGym]] — AI SRE エージェント向けの高忠実度ライブベンチマーク兼フレームワーク
- [[Stratus]] — マルチエージェントの SRE エージェント(4 エージェント + 状態機械、TNR で安全な巻き戻しと再試行。CrewAI 基盤)
- [[CrewAI]] — STRATUS の実装基盤となった LLM マルチエージェントフレームワーク
- [[AIOpsLab]] — 自律クラウド向け AIOps エージェント評価フレームワーク(AgentOps を提唱、SREGym の先行ベンチマーク)
- [[ITBench]] — SRE・CISO・FinOps の 3 ペルソナ横断で AI エージェントを実環境評価する IT 自動化ベンチマーク(IBM/UIUC、ICML'25。一次論文 ingest 済み)
- [[ChaosMesh]] — Kubernetes 向けカオスエンジニアリング / 障害注入ツール(AIOpsLab が symptomatic fault の注入に統合)
- [[PAGER]] — エンタープライズ AI アシスタント向けのプロアクティブな障害予測・説明・対話支援エージェント(Adobe)
- [[Adobe Experience Platform]] — Adobe の大規模カスタマーデータプラットフォーム(PAGER の対象運用環境)
- [[MicroRemed]] — エンドツーエンドのマイクロサービス修復を評価する初のライブベンチマーク(PKU/Alibaba)
- [[ThinkRemed]] — SRE 的な反復推論を模すマイクロサービス修復のマルチエージェントフレームワーク(MicroRemed の参照手法)
- [[Ansible]] — 宣言的・エージェントレスな IT 自動化フレームワーク(MicroRemed の緩和アクション出力形式)
- [[AI Operator]] — Google の自律的な一次対応エージェント(並列調査・L2/L3 稼働、推論を [[Actus]] と分離)
- [[Actus]] — Google のアクチュエーション統一コントロールプレーン兼セーフティゲートウェイ(dry-run・"Red Button")
- [[Detectr]] — Google の Gemini 駆動の障害検知プラットフォーム(ユーザーフィードバックを一次シグナル化)
- [[Model Context Protocol]] — AI エージェントとツール接続を標準化するオープン仕様(MCP、Production Agent/A2A)
- [[Bits AI SRE]] — Datadog の自律インシデント調査・RCA エージェント(仮説駆動、調査・RCA 段に特化)
- [[Toto]] — Datadog の観測データ特化のゼロショット時系列予測基盤モデル(151M、decoder-only)
- [[Falcon-X]] — Ant International の異種多変量向けの encoder-only 時系列基盤モデル(591M、潜在プロトタイプルーティング)
- [[Chronos-2]] — group attention で多変量・文脈内学習を可能にした時系列基盤モデル(Falcon-X の主要比較対象)
- [[MetricSifter]] — 障害箇所特定の前処理となる特徴削減フレームワーク(変化点検知 + KDE、教師なし・多変量、SAKURA Internet)
- [[PyRCA]] — メトリクスベースの根本原因分析ライブラリ(MetricSifter の合成データ生成器兼 FL ベースライン、Salesforce)
- [[HeteroTSDB]] — 異種 KVS を TTL ベースで階層 tiering する TSDA(memory-KVS + disk-KVS、KairosDB 比 3.98 倍の ingestion、Mackerel に実投入)
- [[go-conntracer-bpf]] — カーネル内のフローバンドリングを実装した eBPF ネットワークフロートレーサの Go ライブラリ(博士論文 Chapter 3 の社会実装)
- [[Mackerel]] — Hatena の SaaS サーバ監視サービス。HeteroTSDB の本番採用先(mackerel.io)
- [[MonitorAssistant]] — LLM(GPT-4 Turbo)ベースのエンドツーエンドの実用的異常検知システム(設定推奨・異常レポート・LLM-Engineer-In-The-Loop、Microsoft 投入)
- [[LogPilot]] — アラート定義(PromQL)の意図を解釈してログを絞り、request-centric な log chain のクラスタリングで根本原因を診断する LLM ベースのアラート診断フレームワーク(CUHK×ByteDance、Volcano Engine に本番展開)
- [[Volcano Engine]] — ByteDance のクラウドプラットフォーム。LogPilot の本番デプロイ先・データ出所で、LLM backend の Doubao シリーズを提供
- [[Toto-1.0-QA-Experimental]] — 事前学習済み TSFM [[Toto]] を VLM([[Qwen3-VL]])と結合した ARFBench 用の時系列 QA モデル。精度 63.9% でフロンティアモデルに匹敵(Datadog/CMU)
- [[Qwen3-VL]] — Alibaba のビジョン言語モデル。Toto-1.0-QA-Experimental のバックボーン
- [[TimesFM]] — Google の decoder-only 時系列基盤モデル。Cisco TSM の継続事前学習のベースモデル
- [[TiRex]] — NXAI の zero-shot 時系列基盤モデル(enhanced in-context learning で長短ホライズン対応、Auer+ 2025)。TimeCopilot の MedianEnsemble 構成要素
- [[Splunk Observability Cloud]] — Splunk の観測プラットフォーム。Cisco TSM の観測時系列データの出所
- [[RFT-FM]] — 強化ファインチューニングの障害を検知→診断→修復する閉ループ障害管理フレームワーク(RFT Failure Management、PKU)
- [[OpenRLHF]] — RLHF/RFT のオープンソース訓練フレームワーク(RFT-FaultBench の障害注入・実行基盤)
- [[R-Pingmesh]] — BUPT/Douyin Vision の能動プロービングに基づくサービス認識型 RoCE ネットワーク監視・診断システム(Agent/Controller/Analyzer、UD QP+CQE タイムスタンプ、ToR-mesh、本番数万 RNIC・6 か月、SIGCOMM 24)
- [[ByteRobust]] — ByteDance の LLM 訓練特化 GPU インフラ管理・障害許容システム(Control/Data Plane、ETTR 最大 97%、迅速な隔離=過剰排除、SOSP 25)
- [[SMon]] — ByteDance×NYU の LLM 学習ストラグラー監視システム。What-if 分析をオンライン化しヒートマップで根本原因を判別(OSDI 25)
- [[NDTimeline]] — ByteDance 内製プロファイラ(veScale 配下)。SMon の What-if 分析の入力トレースを生成、10% サンプリング
- [[Astral]] — Nanjing/Tencent の 50 万 GPU 級 LLM 訓練データセンターインフラ(tier-2 同一レール相互接続+HVDC+空気液体冷却+4 層フルスタック監視+Seer、SIGCOMM 25)
- [[Seer]] — Astral のオペレータ粒度予測コンポーネント。実測スループットで自己補正し数秒でタイムライン生成、密モデル 0.3% 偏差
- [[Delta]] — NCSA 運用の大規模 GPU HPC システム(A100 448+H100/GH200 608、計 1,056 GPU)。GPU レジリエンス研究の対象・データ源
- [[StabilityDB]] — Intel/RIKEN の [[Aurora]] 向け自動故障管理システム。相関イベント分析の集中型メタデータベース + マルチストライク修復ポリシーで MTTR を手動比最大 84 倍短縮(SC 25)
- [[Aurora]] — Argonne National Laboratory の 63,744 GPU エクサスケールスパコン。StabilityDB の実導入先・評価環境
- [[Guard]] — Amazon Store Foundational AI のストラグラー検知 + ノード健全性管理システム。オンラインモニタリング + オフラインノードスイープの閉ループでグレーノードを検知、MFU 最大 1.7 倍(MLSys 26)
- [[fkat]] — Amazon の AI 訓練向け OSS キット。Guard の検知ツールの一部を公開(github.com/amzn/fkat)
- [[FlashRecovery]] — iFLYTEK/USTC/Huawei の LLM 訓練高速障害復旧システム。チェックポイントフリー 1 ステップ復旧 + スケール非依存タスク再起動、4,800 デバイス 150 秒(arXiv:2509.03047)
- [[Ascend NPU]] — Huawei の AI アクセラレータ。FlashRecovery の評価環境(4,800 デバイス)
- [[Hawkeye]] — Tsinghua/Beihang/Infrawaves の RDMA 性能異常(NPA)診断システム。PFC プロベナンスで backpressure/storm/deadlock を 90% 以上の精度で診断(SIGCOMM 25)
- [[Intel Tofino]] — Intel の P4 プログラマブルスイッチ ASIC。Hawkeye のデータプレーン内 PFC 因果関係解析のテストベッド
- [[C4]] — Alibaba/HKUST の通信駆動型大規模 AI 訓練効率化ソリューション(集団通信較正)。診断 [[C4D]] + 性能 [[C4P]] でダウンタイム 31.19%→1.16%・システム効率 30%→45%、本番 30 か月超(HPCA 25)
- [[C4D]] — [[C4]] の診断サブシステム([[ACCL]] 拡張)。集団通信の症候から故障コンポーネントを数十秒で特定・隔離・再開
- [[C4P]] — [[C4]] の性能サブシステム。予測可能な長寿命フローのトラフィック工学で帯域競合を削減
- [[ACCL]] — Alibaba Collective Communication Library。C4D が拡張する集団通信ライブラリ
- [[H800]] — NVIDIA GPU。[[C4]] の評価環境
- [[OptProphet]] — Nankai University の光トランシーバー故障の予測 + 分類フレームワーク。特徴量集約で予測 F1 0.884(平均 1.11 日前)・分類 F1 0.855(APNet 25)
- [[PACE]] — ISAV の因果探索フレームワーク(Pattern and Causal Exploration)。相関クラスタリング + ラグ考慮型 Granger 因果性で ORNL Summit の冷却テレメトリ 7 年から有向因果パスを抽出(ISAV 25)
- [[SkeletonHunter]] — Alibaba Cloud のコンテナ訓練ネットワーク障害診断システム。トラフィックスケルトン推論で probing を 2 桁削減、precision 98.2%/recall 99.3%(SIGCOMM 25)
- [[eACGM]] — eBPF + libnvml + GMM のフルスタック非侵入 ML 監視(CUDA/Python/PyTorch/NCCL/GPU)。GMM 異常検知が 6 ベースラインを上回る(IWQoS 25、OSS shady1543/eACGM)
- [[XPUTimer]] — Alibaba/Ant の発散 LLM 訓練異常診断。非侵入 CPython 計装 + CUDA-GDB の intra-kernel inspecting で O(1) 通信ハング箇所特定。arXiv v2 で [[Flare]] に改名(arXiv 25)
- [[LLMPrism]] — Huawei Cloud のブラックボックス性能診断。スイッチ層 RoCE ネットワークフローのみから並列化戦略・ステップタイムラインを逆推定、[[Platform-X]] で稼働(DSN 25)
- [[L4]] — 大規模 LLM 訓練障害の自動ログ解析診断。cross-job/spatial(Isolation Forest)/temporal(DTW)の 3 パターン、F1 0.873、[[Platform-X]] に展開(ESEC/FSE 25)
- [[Platform-X]] — Huawei Cloud の本番マルチテナント LLM 訓練/開発基盤(論文では匿名)。[[L4]]・[[LLMPrism]] 双方のデプロイ先
- [[Summit]] — Oak Ridge National Laboratory のスパコン。冷却インフラの 7 年分テレメトリ(Yokogawa SMARTDAC)が [[PACE]] の評価データ
- [[Perlmutter]] — NERSC のスパコン(AMD EPYC + A100-SXM4、NVLink3.0、最大 128 A100)。GPU 性能モデリングの評価テストベッド、平均誤差 4.98%
- [[Vista]] — TACC のスパコン(Grace + GH200、NDR InfiniBand、最大 128 GH200)。GPU 性能モデリングの評価テストベッド、平均誤差 9.38%
- [[GPT-NeoX]] — [[DeepSpeed]] + [[Megatron-LM]] を統合し 3D parallelism を実現する訓練フレームワーク。GPU 性能モデリングのオペレータ分解対象
- [[Alibaba HPN]] — Alibaba の rail-optimized な LLM 訓練向け DC ネットワーク(Qian+, SIGCOMM 24)。[[SkeletonHunter]] の前提ネットワーク
- [[DeepSpeed]] — Microsoft のメモリ最適化(ZeRO)+ 3D parallelism 実装基盤([[GPT-NeoX]]・[[Aegis]] が参照)
- [[DyTwin]] — データセンター向けフェデレーテッド適応型デジタルツイン枠組み(HPE/ORNL、SC24-W)。[[PACE]] が因果グラフを注入する統合先
- [[FLASH]] — 反復インシデント診断を自動化する [[Microsoft]] の LLM ワークフローエージェント(status supervision + hindsight、本番 250 件で [[TaskWeaver]] 比 +13.2%)(product / aiops)
- [[StepFly]] — [[TSG自動化]] のエンドツーエンドエージェント型フレームワーク([[Microsoft]]/[[Tsinghua University]]、[[TSG Mentor]] + オフライン DAG+QPP + 並列 scheduler-executor、GPT-4.1 約 94%・実行時間 32.9〜70.4% 削減)(product / aiops)
- [[TSG Mentor]] — [[StepFly]] 第 1 段の TSG 品質改善ツール(品質問題 CP/CF/DF/DI/PS を検知、F1 0.81)(product / aiops)
- [[LLexus]] — TSG を計画フェーズ前置でコンパイルし決定論的に実行する [[Microsoft]] のインシデント管理エージェント([[Semantic Kernel]]+GPT-4-Turbo で計画、[[Azure Durable Functions]] で実行、計画 1 TSG あたり $0.60〜$1.71)(product / aiops)
- [[TaskWeaver]] — [[Microsoft]] のコードファースト LLM エージェントフレームワーク([[FLASH]] の主要ベースライン、FLASH 評価で 60.7%)(product)
- [[Semantic Kernel]] — [[Microsoft]] の LLM オーケストレーション OSS([[LLexus]] の計画生成フレームワーク)(product)
- [[Azure Durable Functions]] — [[Microsoft]] Azure のステートフルなサーバレス実行基盤([[LLexus]] エグゼキュータの決定論的実行基盤)(product)
- [[PMF]] — [[IBM Research]] のプログラマブルメトリクスフローフレームワーク。LP 最適化で collect-first→use-first パラダイム(product / observability)
- [[LogReducer]] — [[Tencent]]/HUST のカーネルログホットスポット動的削減ツール。eBPF + EMFP で [[WeChat]] 本番稼働(product / log analysis)
- [[WeChat]] — [[Tencent]] のメッセージングプラットフォーム。LogReducer の評価環境(product)
- [[LogCleaner]] — ログイベント削減による異常検知モデル強化手法。TF-IDF/クラスタリング/エントロピーの 3 戦略(product / log analysis)
- [[Hindsight]] — [[Max Planck Institute for Software Systems]] の遡及的分散トレースサンプリングシステム。障害後完全トレース収集(product / distributed tracing)
- [[OpenTelemetry]] — テレメトリの標準フレームワーク。Tracezip・Astraea・Mint・TraStrainer の計装基盤(product / observability)
- [[Tracezip]] — [[Sun Yat-sen University]] のトレース圧縮。共通性・変動性分解で圧縮率 80% 超(product / distributed tracing)
- [[Astraea]] — [[Boston University]] のオンライン確率的トレーシング。スパンレベル重要度スコアリング(product / distributed tracing)
- [[VAIF]] — Variable Attention and Influence Function。[[Astraea]] のスパン重要度推定モジュール(product)
- [[Mint]] — [[Sun Yat-sen University]]/[[Alibaba Group]] の全リクエスト収集トレーシング。共通性・変動性分析(product / distributed tracing)
- [[TraStrainer]] — [[Huawei Technologies]] の適応的トレースサンプリング。実行時状態で tail-based を強化(product / distributed tracing)
- [[ByteSeries]] — [[ByteDance]] / [[Huazhong University of Science and Technology]] の本番監視向けインメモリ TSDB。Compressed Inverted Index(trie + p4nzenc64)+ 3 段メモリ構造で tsdc 比メタデータ −60%・多次元クエリ 1.8〜10.7 倍(SoCC 2020)
- [[tsdc]] — [[ByteDance]] の元本番インメモリ TSDB。Gorilla 類似構造だがメタデータがメモリ 80% 超を占め [[ByteSeries]] に置き換えられた(product / time-series)
## Repository / Dataset
- [[OpenTracing Processor]] — Bento+ 2021(J Grid Computing)が公開した OpenTracing データ用処理パイプライン(Java Streaming API で span→trace 再構築、NetworkX で graph 処理、OpenTSDB+Grafana にメトリクス出力)。GitHub: andrepbento/OpenTracingProcessor(repository / observability / tracing)
- [[btree-cpp]] — B-Trees Are Back の unsynchronized B-Tree 実装(repository / database)
- [[btree24]] — B-Trees Are Back の [[vmcache]] 統合版 B-Tree 実装(repository / database)
- [[philly-traces]] — Microsoft [[Philly]] の DNN 訓練クラスタスケジューラトレース。ジョブ到着・サイズ・配置・実行時間を含む
- [[TelecomTS]] — 5G 通信ネットワーク由来の大規模マルチモーダルオブザーバビリティデータセット(非匿名化・絶対スケール保持、32K×18 KPI、221 万 Q&A、Yale University)
- [[DeathStarBench]] — マイクロサービスのベンチマークスイート(AIOpsLab・SREGym 双方のテストベッド)
- [[Train-Ticket]] — 鉄道予約題材の大規模マイクロサービスベンチマーク(MicroRemed の最難環境、Zhou+ 2018)。[[RCAEval]] 2025 では 64 サービスの最大ベンチマーク用マイクロサービスとして採用
- [[Online-Boutique]] — Google のクラウドファーストなマイクロサービスのデモ(= microservices-demo、MicroRemed 採用)。[[RCAEval]] 2025 では 12 サービスの gRPC ベース e コマースとして採用
- [[Sock Shop]] — 靴下販売題材のマイクロサービスベンチマーク(MetricSifter の実証研究、7 microservices、Weaveworks)。[[RCAEval]] 2025 では 15 サービスの HTTP ベンチマークとして採用
- [[RCAEval]] — マイクロサービス RCA の評価フレームワーク兼ベンチマーク([[RMIT University]] / [[University of Newcastle]] / [[University of New South Wales]])。RE1/RE2/RE3 合計 735 ケース・11 種障害(リソース 4 + ネットワーク 2 + コードレベル 5)、15 ベースラインを統一フレームワークで評価可能(repository / aiops / benchmark)
- [[Meltria]] — マイクロサービスの障害データセット生成基盤(MetricSifter の実証データ作成、github.com/ai4sre/meltria)
- [[BOOM]] — 実運用テレメトリのみで構成した観測時系列予測ベンチマーク(2,807系列・約3.5億点、Datadog)
- [[GIFT-Eval]] — 汎用時系列予測モデル評価のベンチマーク(15 univariate + 8 multivariate、7 ドメイン、144K 系列)
- [[TIME]] — 汚染耐性の zero-shot 時系列予測ベンチマーク(Qiao et al. 2026)。公開データセットへの汚染を防ぐために新規収集されたデータで構成(dataset / time-series)
- [[fev-bench]] — 現実的な時系列予測ベンチマーク(100 タスク、観測系 BOOMLET= BOOM 部分集合を含む)
- [[ARFBench]] — ソフトウェアインシデント対応の時系列質問応答(TSQA)を測る初のベンチマーク(750 問・142 系列・538 万点、Datadog 本番インシデントの Slack タイムライン由来)
- [[RFT-FaultBench]] — 強化ファインチューニングの初の細粒度障害ベンチマーク(5 families/16 types/779 runs/22,549 train-step/145 万 trajectory、PKU)
- [[StragglerAnalysis]] — ストラグラー分析論文の公式 artifact リポジトリ(ByteDance-Seed/StragglerAnalysis、Python 3.11)
- [[DLRover]] — Ant Group の OSS 自動分散 DL システム(LF AI & Data Foundation 支援)。[[Flare]]([[XPUTimer]])をコア部品として含む
- [[The Pile]] — 825GB の多様な英語テキストコーパス。GPU 性能モデリングの訓練・評価コーパス([[GPT-NeoX]]-20B トークナイザ)
- [[LLM4Log (repository)]] — LLM4Log サーベイのコンパニオンリポジトリ(github.com/zeyang919/LLM4Log)。145 論文を 7 タスクで分類した論文リスト + コーパスメタデータ + 図(Concordia SPEAR lab)
## Person
- [[Xupeng Miao]] — LLM サービングサーベイ筆頭著者(Purdue University、ACM Computing Surveys 2025)
- [[Zhihao Jia]] — CMU 助教授。FlexFlow・SpecInfer の創始者、LLM サービングサーベイ責任著者
- [[Tianqi Chen]] — CMU 教授。TVM・XGBoost・MLC-LLM の創始者、MLSys 分野の著名研究者
- [[Jeffrey C. Mogul]] — Google Research。可用性・SLO 研究の主要著者(HotOS 2017・HotOS 2019)
- [[John Wilkes]] — Google。SLO/SRE の著名研究者(HotOS 2019、SRE 本共著者)
- [[Tamás Hauer]] — Google。ウィンドウ付きユーザーアップタイムの提案者(NSDI 2020 筆頭著者)
- [[Philipp Hoffmann]] — Google。Meaningful Availability 共著者(NSDI 2020)
- [[John Lunney]] — Google。Meaningful Availability 共著者(NSDI 2020)
- [[Dan Ardelean]] — Google。Meaningful Availability 共著者(NSDI 2020)
- [[Amer Diwan]] — Google。Meaningful Availability 共著者(NSDI 2020)
- [[Rebecca Isaacs]] — Google。可用性のセキュリティ的思考を提唱(HotOS 2017 共著者)
- [[Brent Welch]] — Google。可用性のセキュリティ的思考を提唱(HotOS 2017 共著者)
- [[Boris Sedlak]] — TU Wien。SLO 拡散方法論の筆頭著者(IEEE SOSE 2024)
- [[Víctor Casamayor Pujol]] — TU Wien。SLO 拡散方法論の共著者(IEEE SOSE 2024)
- [[Praveen Kumar Donta]] — TU Wien。SLO 拡散方法論の共著者(IEEE SOSE 2024)
- [[Juan Luis Herrera]] — TU Wien。CASCA 筆頭著者(arXiv 2026)
- [[Daniel Wang (TU Wien)]] — TU Wien。CASCA 共著者(arXiv 2026)
- [[Zeyang Ma]] — LLM4Log サーベイ筆頭著者(Concordia University SPEAR lab、ログ解析)
- [[Jinqiu Yang]] — LLM4Log サーベイ共著者(Concordia University、ソフトウェア工学)
- [[Tse-Hsun Chen]] — LLM4Log サーベイ senior 著者・SPEAR lab 主宰(Concordia University、ログ解析の主要研究拠点の 1 つ)
- [[Martin Casado]] — クラウド管理ビジョン論文の共著者(Andreessen Horowitz、GoEx 共著)
- [[Archit Bhatnagar]] — クラウド管理ビジョン論文の共同筆頭著者(University of Michigan)
- [[Tongyuan Miao]] — クラウド管理ビジョン論文の共著者(University of Michigan)
- [[Yunming Xiao]] — クラウド管理ビジョン論文の共著者(University of Michigan、Pulse の Yibo Xiao とは別人)
- [[Yibo Huang]] — クラウド管理ビジョン論文の共著者・cloudless computing 共著(University of Michigan)
- [[Ali Maatouk]] — TelecomTS の corresponding author(Yale University)
- [[Rex Ying]] — TelecomTS の senior author(Yale University)
- [[Tianyin Xu]] — SREGym の最終著者(UIUC)
- [[Yinfang Chen]] — AIOpsLab・SREGym・STRATUS の第一/共著者(UIUC)
- [[Saurabh Jha]] — ITBench 主導著者・STRATUS 共著者(IBM Research)
- [[Rohan Arora]] — ITBench equal-contribution リード・STRATUS 共著者(IBM Research)
- [[Minghua Ma]] — AIOpsLab corresponding author(Microsoft、AIOps 研究を多数主導)
- [[Yunyao Li]] — PAGER のシニア著者(Adobe、enterprise AI assistant 研究)
- [[Lingzhe Zhang]] — MicroRemed 等の第一著者(PKU、AIOps/LLM for SRE を多作)
- [[Tong Jia]] — MicroRemed corresponding author(PKU、AIOps 研究を主導)
- [[Ameet Talwalkar]] — Toto 論文の senior author(CMU 兼 Datadog)
- [[Yuuki Tsubouchi]] — MetricSifter 筆頭著者・博士論文著者・SAKURAONE 共著者・本 vault 所有者(SAKURA Internet Research Center、元 Hatena SRE)
- [[Hirofumi Tsuruta]] — MetricSifter 第 2 著者・SAKURAONE 共著者(SAKURA Internet Research Center、機械学習)
- [[Fumikazu Konishi]] — SAKURAONE 論文の筆頭著者・corresponding author(SAKURA Internet Research Center)
- [[Ryosuke Matsumoto]] — Transtracer / socket-based tracing 論文の共著者(博士論文 Chapter 3 の基)
- [[Jiangfei Duan]] — LLM 訓練システムサーベイの筆頭著者(Shanghai AI Lab / CUHK)
- [[Peng Sun]] — LLM 訓練システムサーベイの corresponding author(Shanghai AI Lab)。[[Acme]] ワークロード特徴づけ記事の著者
- [[Ziheng Jiang]] — MegaScale 論文の筆頭著者(ByteDance、equal contribution)
- [[Xin Jin]] — MegaScale 論文の責任著者(Peking University)
- [[Xin Liu]] — MegaScale 論文の責任著者(ByteDance)
- [[Yangtao Deng]] — Minder 論文の筆頭著者(Tsinghua University、ByteDance との共同研究)
- [[Zhuo Jiang]] — Minder 論文の責任著者(ByteDance High-speed Network チーム)
- [[Minlan Yu]] — Minder 論文の責任著者(Harvard University、ネットワーキングの計測・診断)
- [[Yibo Xiao]] — Pulse 論文の筆頭著者(Nanjing University)
- [[Qingkai Meng]] — Pulse 論文の責任著者(Nanjing University、Astral 筆頭著者)
- [[Chen Tian]] — Pulse 論文のシニア著者(Nanjing University、μMon/Astral のネットワーク計測系譜)
- [[Zhaoyang Yu]] — MonitorAssistant の第一著者(Tsinghua University & BNRist、異常検知・RCA に注力)
- [[Yu Luo]] — OpsAgent 筆頭著者(南開大学 College of Software)。マイクロサービス IM 向け自己進化型 MAS を設計
- [[Dan Pei]] — Tsinghua University & BNRist の教授。NetManAIOps グループを主宰(AIOps・異常検知・RCA)
- [[Changhua Pei]] — CNIC/CAS の研究者(Dan Pei とは別人)。Flow-of-Action の第一著者でマイクロサービス RCA・AIOps を研究
- [[Bowen Hao]] — [[Nankai University]] 所属。CMDiagnostor(WWW 2023)の共著者(person / aiops / rca)
- [[Mingjie Li]] — [[Tsinghua University]]/BNRist 所属。CMDiagnostor(WWW 2023)の共著者(person / aiops / rca)
- [[Xianglin Lu]] — [[Tsinghua University]]/BNRist 所属。CMDiagnostor(WWW 2023)の共著者(person / aiops / rca)
- [[Haoran Yan]] — GenAI クラウドインシデント実証研究の第一著者(HUST、ICSE 2026)
- [[Ying Li]] — LLM4AIOps サーベイの責任著者(Peking University。Lingzhe Zhang・Tong Jia と PKU の AIOps クラスタを形成)
- [[Philip S. Yu]] — LLM4AIOps サーベイ共著者(University of Illinois Chicago。データマイニング/機械学習の著名研究者)
- [[Aoyang Fang]] — 障害伝播を意識した RCA ベンチマーク論文の第一著者(CUHK-Shenzhen)
- [[Pinjia He]] — 障害伝播を意識した RCA ベンチマーク論文の corresponding author(CUHK-Shenzhen、マイクロサービス RCA・障害診断)
- [[Zhihan Jiang]] — LogPilot 筆頭著者(CUHK、Michael R. Lyu グループ)。ログ解析・ログ parsing(LILAC)・障害診断を専門
- [[Michael R. Lyu]] — CUHK 教授。ソフトウェア信頼性・ログ解析・AIOps の著名研究者で、Drain・LILAC・COCA など多数のログ解析研究と LogPilot を主導(ログ解析の二大ハブの一つ、もう一方は [[Dan Pei]])
- [[Tieying Zhang]] — LogPilot の corresponding author(ByteDance)。Volcano Engine への本番展開・データ収集を取りまとめ
- [[Stephan Xie]] — ARFBench 第一著者([[Carnegie Mellon University]] / [[Datadog]])
- [[Ben Cohen]] — ARFBench 共著者([[Datadog]] AI Research、Toto 系列の開発に関与)
- [[Mononito Goswami]] — ARFBench 共著者([[Carnegie Mellon University]] / [[Amazon Web Services]])
- [[Liang Gou]] — Cisco Time Series Model の著者([[Cisco]] / [[Splunk]])
- [[Yunpeng Zhai]] — RFT-FM 論文の共著者(PKU/Alibaba/UIC グループ)
- [[Liancheng Fang]] — RFT-FM 論文の共著者([[University of Illinois Chicago]])
- [[Kening Zheng]] — RFT-FM 論文の共著者([[Peking University]])
- [[Hongyi Liu]] — RFT-FM 論文の共著者(PKU/Alibaba/UIC グループ)
- [[Xiaosong Huang]] — RFT-FM 論文の共著者([[Peking University]])
- [[Siva Rama Krishna Kottapalli]] — TSFM サーベイの筆頭著者([[Dell Technologies]]、6 次元タクソノミーを提案)
- [[Azul Garza]] — [[TimeCopilot]] 筆頭著者。TimeGPT-1 や Nixtla の OSS 予測ツール群(StatsForecast/NeuralForecast/tsfeatures)の著者(所属明記なし、San Francisco)
- [[Renée Rosillo]] — [[TimeCopilot]] 共著者(所属明記なし、San Francisco)
- [[Kefei Liu]] — [[R-Pingmesh]] 筆頭著者([[BUPT]]、Douyin Vision との共同研究)。Hostping(NSDI 2023)筆頭著者でもある(person / networking)
- [[Jiao Zhang]] — [[R-Pingmesh]] 責任著者([[BUPT]] / Purple Mountain Laboratories)(person / networking)
- [[Shengkun Cui]] — [[Delta]] の GPU レジリエンス論文の共同筆頭著者([[University of Illinois Urbana-Champaign]])(person / hpc)
- [[Ravishankar K. Iyer]] — GPU レジリエンス論文の責任著者([[University of Illinois Urbana-Champaign]]、ディペンダブルコンピューティング)(person / hpc)
- [[Jinkun Lin]] — ストラグラー分析(OSDI 25)の筆頭著者([[New York University]])(person / mlsys)
- [[Aurojit Panda]] — ストラグラー分析の共著者([[New York University]])(person / systems)
- [[Jinyang Li]] — ストラグラー分析の責任著者([[New York University]])(person / systems)
- [[Borui Wan]] — [[ByteRobust]] 筆頭著者(equal contribution、[[The University of Hong Kong]]/[[ByteDance]])。ByteCheckpoint 筆頭著者(person / mlsys)
- [[Liang Xiang]] — [[ByteRobust]] の責任著者格([[ByteDance]] Seed)(person / mlsys)
- [[Chuan Wu]] — [[ByteRobust]] 責任著者([[The University of Hong Kong]])(person / mlsys)
- [[Hao Zheng]] — [[Astral]] 共著者([[Nanjing University]])(person / networking)
- [[ChonLam Lao]] — [[Astral]] 共著者([[Harvard University]])(person / networking)
- [[Gianni Antichi]] — [[Astral]] 共著者(Politecnico di Milano / Queen Mary University of London、プログラマブルネットワーク)(person / networking)
- [[Pavana Prakash]] — [[PACE]](ISAV)筆頭著者([[Hewlett Packard Labs]])(person / reliability)
- [[Rolando P. Hong Enriquez]] — [[PACE]] 共著者([[Hewlett Packard Labs]])(person)
- [[Sergey Serebryakov]] — [[PACE]] 共著者([[Hewlett Packard Labs]])(person)
- [[David Grant]] — [[PACE]] 共著者(所属は一次ソース未確定)(person)
- [[Wesley Brewer]] — [[PACE]] 共著者([[Oak Ridge National Laboratory]]、Summit 評価データ側)(person / hpc)
- [[Dejan Milojicic]] — [[PACE]] シニア著者([[Hewlett Packard Labs]]、HPE Fellow/VP)(person)
- [[Biyao Zhang]] — GPU 性能モデリング論文の筆頭著者([[Case Western Reserve University]])(person / mlsys)
- [[Mingkai Zheng]] — 同共著者([[Rutgers University]])(person)
- [[Debargha Ganguly]] — 同共著者([[Case Western Reserve University]])(person)
- [[Xuecen Zhang]] — 同共著者([[Case Western Reserve University]])(person)
- [[Vikash Singh]] — 同共著者([[Case Western Reserve University]])(person)
- [[Vipin Chaudhary]] — 同共著者([[Case Western Reserve University]])(person)
- [[Zhao Zhang]] — 同シニア著者([[Rutgers University]]、HPC)(person / hpc)
- [[Wei Liu]] — [[SkeletonHunter]] 筆頭著者([[Tsinghua University]])(person / networking)
- [[Kun Qian]] — [[SkeletonHunter]]/[[Aegis]] 共著者・[[Alibaba HPN]] 筆頭(Alibaba Cloud)(person / networking)
- [[Zhenhua Li]] — [[SkeletonHunter]] 共著者([[Tsinghua University]])(person)
- [[Ennan Zhai]] — [[SkeletonHunter]]/[[Aegis]] の責任著者格(Alibaba Cloud、DC ネットワーク)(person / networking)
- [[Yunhao Liu]] — [[SkeletonHunter]] 共著者([[Tsinghua University]])(person)
- [[Weicheng Wang]] — [[SkeletonHunter]]/[[Aegis]] 共著者(Alibaba Cloud)(person)
- [[Yun Zhang]] — [[SkeletonHunter]] 共著者(Alibaba Cloud)(person)
- [[Jiakang Li]] — [[SkeletonHunter]] 共著者(Alibaba Cloud)(person)
- [[Shuhong Zhu]] — [[SkeletonHunter]] 共著者(Alibaba Cloud)(person)
- [[Xue Li]] — [[SkeletonHunter]]/[[Aegis]] 共著者(Alibaba Cloud)(person)
- [[Hongfei Xu]] — [[SkeletonHunter]] 共著者(Alibaba Cloud)(person)
- [[Fei Feng]] — [[SkeletonHunter]]/[[Aegis]] 共著者(Alibaba Cloud)(person)
- [[Ruilin Xu]] — [[eACGM]] 筆頭著者([[Sun Yat-sen University]])(person)
- [[Zongxuan Xie]] — [[eACGM]] 共著者([[Sun Yat-sen University]])(person)
- [[Weihao Cui]] — [[XPUTimer]]([[Flare]])筆頭著者([[Shanghai Jiao Tong University]])(person / mlsys)
- [[Ji Zhang]] — 同共著者(v2 所属 Independent Researcher)(person)
- [[Han Zhao]] — 同 corresponding(v2)([[Shanghai Jiao Tong University]])(person)
- [[Chao Liu]] — 同共著者(v2 所属 Independent Researcher)(person)
- [[Jian Sha]] — 同 corresponding(v2)([[Ant Group]])(person)
- [[Bingsheng He]] — 同共著者([[National University of Singapore]])(person)
- [[Xuanhua Shi]] — [[Huazhong University of Science and Technology]] 教授。インメモリ TSDB [[ByteSeries]](SoCC 2020)の筆頭著者(person / database / time-series)
- [[Yongluan Zhou]] — [[University of Copenhagen]] 所属。[[ByteSeries]](SoCC 2020)の共著者(person / database)
- [[Minyi Guo]] — 同シニア著者([[Shanghai Jiao Tong University]])(person)
- [[Quan Chen]] — 同 corresponding(v1)([[Shanghai Jiao Tong University]])(person)
- [[Jianbo Dong]] — [[Aegis]] 筆頭著者(equal contribution、Alibaba Cloud)(person)
- [[Pengcheng Zhang]] — [[Aegis]] 共著者(Alibaba Cloud)(person)
- [[Rui Ren]] — [[LLMPrism]] 共著者([[Huawei Cloud]])(person)
- [[Yulun Wu]] — [[LLMPrism]] 共著者([[The Chinese University of Hong Kong]])(person)
- [[Wenwei Gu]] — [[LLMPrism]] 共著者([[The Chinese University of Hong Kong]])(person)
- [[Yujie Huang]] — [[LLMPrism]] 共著者([[The Chinese University of Hong Kong]])(person)
- [[Junjie Huang]] — [[L4]] 筆頭著者([[The Chinese University of Hong Kong]])(person)
- [[Zhuangbin Chen]] — [[L4]] 共著者([[Sun Yat-sen University]])(person)
- [[Yichen Li]] — [[L4]]/[[LLMPrism]] 共著者([[The Chinese University of Hong Kong]])(person)
- [[Renyi Zhong]] — [[L4]] 共著者([[The Chinese University of Hong Kong]])(person)
- [[Cong Feng]] — [[L4]]/[[LLMPrism]] 共著者([[Huawei Cloud]])(person)
- [[Yongqiang Yang]] — [[L4]]/[[LLMPrism]] 共著者([[Huawei Cloud]])(person)
- [[Zengyin Yang]] — [[L4]]/[[LLMPrism]] 共著者([[Huawei Cloud]])(person)
- [[Xuchao Zhang]] — [[FLASH]] 筆頭著者([[Microsoft]] Redmond)(person / aiops)
- [[Tanish Mittal]] — [[FLASH]] 共著者([[Microsoft]] Bengaluru)(person)
- [[Chetan Bansal]] — [[FLASH]] 共著者([[Microsoft]] Redmond、クラウドインシデント RCA 研究)(person / aiops)
- [[Rujia Wang]] — [[FLASH]] 共著者([[Microsoft]] Redmond)(person)
- [[Zhixin Ren]] — [[FLASH]] 共著者([[Microsoft]] Redmond)(person)
- [[Hao Huang]] — [[FLASH]] 共著者([[Microsoft]] Redmond)(person)
- [[Saravan Rajmohan]] — [[FLASH]]・[[StepFly]] 双方の共著者([[Microsoft]] Redmond)。[[TSG自動化]] 2 本の結節点(person / aiops)
- [[Jiayi Mao]] — [[StepFly]] 筆頭著者([[Tsinghua University]])(person)
- [[Liqun Li]] — [[StepFly]] 共著者([[Microsoft]])(person)
- [[Yanjie Gao]] — [[StepFly]] 共著者([[Microsoft]] Research / [[Renmin University of China]])(person)
- [[Zegang Peng]] — [[StepFly]] 共著者([[Tsinghua University]])(person)
- [[Si Qin]] — [[StepFly]] 共著者([[Microsoft]])(person)
- [[Samia Khalid]] — [[StepFly]] 共著者([[Microsoft]] USA)(person)
- [[Sitaram Lanka]] — [[StepFly]] 共著者([[Microsoft]] USA)(person)
- [[Dongmei Zhang]] — [[StepFly]] 共著者([[Microsoft]])(person)
- [[Pedro Las-Casas]] — [[LLexus]] 筆頭著者([[Microsoft]])(person)
- [[Alok Kumbhare]] — [[LLexus]] 共著者([[Microsoft]])(person)
- [[Rodrigo Fonseca]] — [[LLexus]] 共著者([[Microsoft]])(person)
- [[Sharad Agarwal]] — [[LLexus]] 共著者([[Microsoft]])(person)
- [[Muhammad Bilal]] — NetOps/AIOps サーベイの責任著者(Lancaster University)(person / networking)
- [[Jon Crowcroft]] — NetOps/AIOps サーベイ共著者(University of Cambridge、ネットワーキングの著名研究者)(person / networking)
- [[Ruizhi Wang]] — NetOps/AIOps サーベイ共著者(Nanjing University of Information Science and Technology)(person)
- [[Xiaolong Xu]] — NetOps/AIOps サーベイ共著者(Nanjing University of Information Science and Technology)(person)
- [[Schahram Dustdar]] — NetOps/AIOps サーベイ共著者(TU Wien + ICREA Barcelona、サービスコンピューティングの著名研究者)(person / networking)
- [[Xuanhe Zhou]] — LLM × DATA サーベイの Co-first author・DBAIOps 共著者(Shanghai Jiao Tong University、データベース × LLM)
- [[Guoliang Li]] — LLM × DATA サーベイのシニア著者・DBAIOps 共著者(Tsinghua University、データベース × AI)
- [[Wei Zhou]] — DBAIOps の筆頭著者(Shanghai Jiao Tong University、データベース O&M 自動化・LLM×DB)
- [[DBAIOps]] — 知識グラフ + 推論 LLM によるデータベース O&M 自動化システム(25 DB 対応・20 実環境稼働)
- [[Baisheng Technology]] — DBAIOps を開発・運営する深圳の企業(百盛(深圳)科技)
- [[Jonathan Mace]] — [[Max Planck Institute for Software Systems]] の研究者。[[Hindsight]] 筆頭著者(person / distributed tracing)
- [[Kangjin Wang]] — [[IBM Research]] の研究者。[[PMF]] 筆頭著者(person / observability)
- [[Zibin Zheng]] — [[Sun Yat-sen University]] 教授。[[Tracezip]] 責任著者(person / distributed systems)
- [[Mehmet Toslali]] — [[Boston University]] の研究者。[[Astraea]] 筆頭著者(person / distributed tracing)
- [[Ayse K. Coskun]] — [[Boston University]] 教授。[[Astraea]] 責任著者(person / distributed tracing)
- [[Haiyu Huang]] — [[Huawei Technologies]] の研究者。[[TraStrainer]] 筆頭著者(person / distributed tracing)
## Organization
- [[Purdue University]] — 米国インディアナ州の研究大学。LLM サービングサーベイ筆頭著者 Xupeng Miao の所属
- [[Concordia University]] — カナダ・モントリオールの大学。SPEAR lab を擁し LLM4Log サーベイを生む。ログ解析の主要研究拠点の 1 つ
- [[University of California, Berkeley]] — クラウド管理ビジョン論文の共同所属(Yiming Qiu の第 2 所属)
- [[Andreessen Horowitz]] — クラウド管理ビジョン論文の産業側共同所属(Martin Casado、a16z)
- [[University of Illinois Urbana-Champaign]] — SRE/AIOps エージェント評価研究の主要機関(SREGym・AIOpsLab)
- [[Microsoft]] — AIOpsLab の主要所属(AIOps/incident management 研究拠点)
- [[Adobe]] — Adobe Experience Platform と PAGER を擁する企業
- [[Peking University]] — MicroRemed の主所属(AIOps/LLM for SRE 研究拠点)
- [[Alibaba Group]] — MicroRemed の共同所属、Qwen3 シリーズの開発元
- [[Lenovo]] — PC・サーバ大手(天津)。OpsAgent を 53 日間・10,492 件の本番環境で展開し有効性を実証した産業パートナー
- [[Google]] — SRE 発祥企業。Bigtable・GFS・Chubby 等の分散基盤と、AI-Ops を Cloud/Ads/YouTube/Search の本番に展開する組織
- [[Datadog]] — オブザーバビリティ/監視 SaaS ベンダ。自律 SRE エージェント Bits AI SRE・時系列基盤モデル Toto・本番接地型コード最適化 DODO を開発(産業界 2 例目)
- [[DODO]] — Datadog Observability-Driven Optimizer。CPU プロファイル+Live Debugger 実呼び出しからベンチマークを生成し LLM エージェントで Go コードを最適化するシステム([[Junaid Ahmed]]・[[Piotr Bejda]])
- [[Junaid Ahmed]] — Datadog AI Research エンジニア。[[DODO]] 共同開発者
- [[Piotr Bejda]] — Datadog AI Research エンジニア。[[DODO]] 共同開発者
- [[Carnegie Mellon University]] — Toto 論文に Datadog AI Research と共同参加した大学
- [[Shanghai AI Laboratory]] — LLM 訓練システムサーベイの主所属。InternLM/InternEvo を擁する中国の AI 研究機関
- [[ByteDance]] — 10,000 GPU 超 AI クラスタで LLM を訓練する企業。MegaScale の開発・運用主体(veScale OSS)
- [[Ant International]] — 時系列基盤モデル Falcon-X を開発した企業組織(連絡先 @ant-intl.com)
- [[SAKURA Internet]] — 日本のクラウド事業者。SAKURA Internet Research Center を擁し MetricSifter を生んだ
- [[Hatena]] — 監視 SaaS Mackerel を運営する日本企業。Yuuki Tsubouchi の元勤務先(2013–2018, SRE)
- [[Kyoto University]] — Yuuki Tsubouchi に博士号(情報学)を授与した大学。本博士論文の発行機関
- [[Tsinghua University]] — Minder 論文の筆頭著者 Yangtao Deng らの所属大学(中国、ByteDance と共同研究)
- [[Harvard University]] — Minder 論文の責任著者 Minlan Yu の所属大学(米国)
- [[IBM Research]] — ITBench と STRATUS の IBM 側拠点(Saurabh Jha ら、UIUC と連携)
- [[Nanjing University]] — Pulse の主所属(State Key Laboratory of Novel Software Technology。ネットワーク計測/RDMA、μMon・Astral を産む)
- [[National University of Singapore]] — Pulse 論文の共著者 Haifeng Sun の所属(シンガポール、Nanjing University と連携)
- [[University of Illinois Chicago]] — 米国シカゴの研究大学(UIC)。LLM4AIOps サーベイ共著者 [[Philip S. Yu]] の所属(既出の UIUC とは別大学)
- [[Yale University]] — 米国コネチカット州の研究大学。TelecomTS の主所属(Ali Maatouk・Rex Ying)
- [[Huazhong University of Science and Technology]] — 中国・武漢の研究大学(HUST)。GenAI クラウドインシデント実証研究の第一著者 Haoran Yan の所属
- [[The Hong Kong University of Science and Technology (Guangzhou)]] — 中国・広州の研究大学(HKUST-GZ)。LLM4AIOps サーベイ共著者 Xuming Hu の所属
- [[The Chinese University of Hong Kong, Shenzhen]] — 中国・深圳の研究大学(CUHK-Shenzhen)。障害伝播を意識した RCA ベンチマーク論文の主所属(香港の CUHK とは別キャンパス)
- [[The Chinese University of Hong Kong]] — 香港の研究大学(CUHK)。Michael R. Lyu グループがログ解析・AIOps を牽引し、LogPilot を ByteDance と共同開発(深圳の CUHK-Shenzhen とは別大学)
- [[Mingyue Cheng]] — USTC の研究者。ATSF ポジションペーパー第 1 著者、公式コード `Mingyue-Cheng/atsf` を管理。時系列×エージェント研究グループの中心
- [[Xiaoyu Tao]] — USTC の研究者。ATSF 第 2 著者で AgenticRL 実装 [[Cast-R1]]・MemCast・TokenCast の筆頭著者
- [[Qi Liu]] — USTC の責任著者(時系列×エージェント。本 wiki の別人 [[Xin Liu]] とは別人)
- [[Enhong Chen]] — USTC のシニア研究者。同グループの時系列・エージェント研究に一貫して参加
- [[University of Science and Technology of China]] — 中国・合肥の大学(USTC)。State Key Laboratory of Cognitive Intelligence。ATSF 著者全員の所属
- [[Cast-R1]] — エージェント型時系列予測の RL 実装(ATSF の AgenticRL パラダイム代表、Tao+ 2026b、arXiv:2602.13802)
- [[TimeCopilot]] — 複数 TSFM と LLM を統一 API 下に集約するオープンソースなエージェント型予測フレームワーク(ATSF の Workflow パラダイム代表、本体論文 [[@2025__arXiv__TimeCopilot]] ingest 済み、GIFT-Eval CRPS で SOTA、Garza & Rosillo 2025、arXiv:2509.00616)
- [[Yusheng Zheng]] — [[eunomia-bpf]] の中心人物。eBPF×AI の研究・OSS を主導し、eBPF×AI 総説の著者(eBPF / observability)
- [[eunomia-bpf]] — eBPF×AI を主題とする OSS コミュニティ。[[bpftime]]/[[GPTtrace]]/[[AgentSight]]/[[Kgent]]/MCPtrace/ActPlane を開発(eBPF / observability)
- [[bpftime]] — ユーザ空間 eBPF ランタイム。uprobe 高速化や eGPU(eBPF バイトコードを GPU へ PTX/SPIR-V 注入)を実現(eBPF / GPU)
- [[GPTtrace]] — 自然言語からカーネルトレース用 eBPF を LLM で生成するツール(AI for eBPF、背後の研究は [[Kgent]]、姉妹に MCPtrace)(eBPF / agent)
- [[AgentSight]] — eBPF でゼロ計装の LLM/AI エージェント可観測性を実現(TLS 傍受 + カーネルシグナル + 二次 LLM 分析、claude code/gemini-cli を <3% オーバーヘッドで観測、arXiv:2508.02736)(eBPF / observability / agent)
- [[Kgent]] — 初の LLM 駆動 eBPF 合成ツール(KEN)。Z3 ベースの記号検査 + テストで約 80% の意味的正しさ(eBPF'24)(eBPF / agent)
- [[NSync]] — IaC reconciliation のための初の自動エージェントシステム([[University of Michigan]]+[[Amazon Web Services]])。API トレースから drift を検知し Terraform 構成へパッチ(IaC / aiops)
- [[Lilac]] — IaC lifting のニューロシンボリックなルール抽出パイプライン([[University of Michigan]]+[[University of California, San Diego]])。順方向デプロイの観測から逆生成ルールを学習(IaC / aiops)
- [[Amazon Web Services]] — クラウドプロバイダ(AWS)。NSync の共同研究機関で、CloudTrail/Bedrock/Boto3/SSM runbook を基盤に提供(cloud / iac)
- [[University of California, San Diego]] — 米国の研究大学(UCSD)。Lilac の共同研究機関(Zheng Guo)(IaC)
- [[Zhenning Yang]] — NSync 第一著者([[University of Michigan]])。Cloud Infrastructure Management in the Age of AI Agents の筆頭でもある(person / iac)
- [[Jingjia Peng]] — Lilac 第一著者([[University of Michigan]])。研究プロトタイプ Lilac-v0 の作者(person / iac)
- [[AWS CloudTrail]] — AWS の API 監査ログサービス。インターフェース不問で全変更を記録し、NSync の drift 検知の観測点(cloud / iac)
- [[aztfexport]] — Microsoft の Azure 特化 IaC lifting ツール(約 95% カバレッジだが単一 provider 限定)。Lilac の主要比較対象(IaC)
- [[Zodiac]] — クラウド IaC のセマンティックチェックを自動マイニング・デプロイ検証するツール([[University of Michigan]]×[[Microsoft]])。Azure/Terraform で 510 チェックを発掘(IaC / program-analysis)
- [[Terraform]] — HashiCorp の宣言的 IaC フレームワーク(市場最有力)。core compiler は cloud-agnostic で provider 固有チェックはプラグイン経由のみ(Zodiac の semantic gap の構造的要因)(product / iac)
- [[Microsoft Azure]] — Microsoft のパブリッククラウド。Zodiac のセマンティックチェック検証対象(52 リソース種別)(product / cloud)
- [[Ang Chen]] — クラウド IaC × LLM エージェント研究の指導著者(UMich)。Zodiac・[[NSync]]・[[Lilac]] を貫く connecting node(person / iac)
- [[Yiming Qiu]] — Zodiac 筆頭著者・Lilac 共著者(UMich)。クラウド構成・ネットワーク検証(person / iac)
- [[Patrick Tser Jern Kon]] — Zodiac・Lilac 共著者(UMich)。グループの preprint 配布ページ cs-pk.com を運営(person / iac)
- [[Ryan Beckett]] — Zodiac 共著者([[Microsoft]])。ネットワーク構成検証の著名研究者(person / networking)
- [[University of Michigan]] — Zodiac・NSync・Lilac の主所属(Ann Arbor)。Ang Chen グループが Cloud IaC × LLM エージェント研究を牽引(organization)
- [[Cisco]] — ネットワーク機器大手。[[Splunk]] を傘下に持ち、観測ドメイン特化の時系列基盤モデル Cisco TSM を開発(organization)
- [[Splunk]] — 観測・セキュリティの SaaS ベンダ([[Cisco]] 傘下)。Cisco TSM の観測データ([[Splunk Observability Cloud]] 由来)とコード(github.com/splunk)を提供(organization)
- [[Dell Technologies]] — TSFM サーベイの主執筆機関(Hopkinton, MA)。自社モデルでなく 6 次元タクソノミーで TSFM フィールドを俯瞰する立場(organization / time-series)
- [[OpenRCA]] — LLM の RCA 能力を測るベンチ/データセット。335 障害・68.5GB テレメトリ、7 goal 定式化、最良 Claude 3.5 で 11.34%・Hard 0.00%(dataset)
- [[Cloud-OpsBench]] — エージェント型 RCA の再現可能ベンチ。452 障害・40 種・[[Kubernetes]] 全スタック、State Snapshot 決定論的デジタルツイン、過程評価 IAC/RAR/ZTDR(dataset)
- [[AlertGuardian]] — [[Tencent]] 本番のアラートライフサイクル管理フレームワーク。denoise(軽量グラフ)→summary(RAG+LLM)→rule refinement(マルチエージェント)、MTTR 156→21 分・94.8% 削減(product)
- [[Kubernetes]] — コンテナオーケストレーション基盤。Cloud-OpsBench が全スタックを障害注入・診断の対象とする(product)
- [[Junjielong Xu]] — [[OpenRCA]] 筆頭著者([[The Chinese University of Hong Kong, Shenzhen]])(person)
- [[Shilin He]] — [[OpenRCA]] 共著・corresponding author([[Microsoft]])(person)
- [[Qingwei Lin]] — [[OpenRCA]] 共著([[Microsoft]])。AIOps の研究者(person)
- [[Chaoyun Zhang]] — [[OpenRCA]] 共著([[Microsoft]])(person)
- [[Guangba Yu]] — [[AlertGuardian]] 筆頭・[[Cloud-OpsBench]] 責任著者。AIOps/RCA(Nezha/MicroRank/ChangeRCA)。所属が SYSU(AlertGuardian)→CUHK(Cloud-OpsBench)で矛盾、両出典保持(person)
- [[Pengfei Chen]] — [[Sun Yat-sen University]] の研究者。マイクロサービス RCA/異常検知を牽引、[[AlertGuardian]]/[[Cloud-OpsBench]] 共著(person)
- [[Sun Yat-sen University]] — 広州の研究大学(SYSU/中山大学)。Pengfei Chen グループが AIOps・マイクロサービス信頼性を牽引(organization)
- [[Tencent]] — [[AlertGuardian]] の本番投入先(論文では Company-X)・共著。大規模クラウドのアラート運用(organization)
- [[TimeSeriesScientist]] — 単変量時系列予測の全工程を 4 エージェント(Curator/Planner/Forecaster/Reporter)で自動化する初の LLM 駆動エージェント型フレームワーク(TSci)。[[エージェント型時系列予測]] の Workflow パラダイム代表、基盤モデル不使用で 21 モデルライブラリ内蔵(repository)
- [[Haokun Zhao]] — [[TimeSeriesScientist]] 論文の筆頭著者(equal contribution、[[Stony Brook University]]/UC San Diego)(person)
- [[Xiang Zhang]] — [[TimeSeriesScientist]] 論文の共著(equal contribution、University of British Columbia)。LLM・マルチモーダル研究。同名の別人物に注意(person)
- [[Jiaqi Wei]] — [[TimeSeriesScientist]] 論文の共著(equal contribution、Zhejiang University)。LLM・RAG(AlignRAG)研究(person)
- [[Chenyu You]] — [[TimeSeriesScientist]] 論文の corresponding author([[Stony Brook University]])。Y-Research-SBU グループ(person)
- [[Stony Brook University]] — 米国の研究大学。[[Haokun Zhao]]・[[Chenyu You]] の所属、Y-Research-SBU グループの拠点(organization)
- [[BUPT]] — 北京郵電大学(Beijing University of Posts and Telecommunications)。[[R-Pingmesh]] の共同研究機関(State Key Laboratory of Networking and Switching Technology、一部 Purple Mountain Laboratories 兼任)(organization / networking)
- [[Douyin Vision]] — [[R-Pingmesh]] を BUPT と共同開発し自社 DML RoCE クラスタに本番展開した企業(ByteDance 傘下ブランド、原名 Douyin Vision Co., Ltd.)(organization)
- [[NCSA]] — National Center for Supercomputing Applications([[University of Illinois Urbana-Champaign]] 内)。[[Delta]]/DeltaAI を運用し GPU レジリエンス研究に現場知見を提供(organization / hpc)
- [[Nokia Bell Labs]] — GPU レジリエンス論文(SC2025)の共著者 Catello Di Martino の所属(organization)
- [[New York University]] — NYU。ストラグラー分析(OSDI 25)の著者陣の所属、ByteDance Seed との産学共同(organization)
- [[The University of Hong Kong]] — HKU。[[ByteRobust]] の責任著者 Chuan Wu・筆頭 Borui Wan が所属(organization)
- [[Zeying Zhu]] — [[University of Maryland]] の研究者。[[PromSketch]]([[@2025__VLDB__Approximation-First Timeseries Monitoring Query At Scale]])の筆頭著者(person)
- [[Jonathan Chamberlain]] — [[Boston University]] の研究者。[[PromSketch]] 論文の共著者(person)
- [[Kenny Wu]] — [[University of Maryland]] の研究者。[[PromSketch]] 論文の共著者(person)
- [[David Starobinski]] — [[Boston University]] の研究者。[[PromSketch]] 論文の共著者(person)
- [[Zaoxing Liu]] — [[University of Maryland]] の研究者、[[PromSketch]] 論文の最終著者。スケッチベース近似クエリ・ネットワーク計測が専門(person)
- [[University of Maryland]] — 米国の州立大学。[[PromSketch]] 研究グループ([[Zeying Zhu]]・[[Kenny Wu]]・[[Zaoxing Liu]])の拠点(organization)
- [[Boston University]] — 米国の私立大学。[[PromSketch]] 論文の [[Jonathan Chamberlain]]・[[David Starobinski]] の所属(organization)
- [[PromSketch]] — 時系列モニタリング向けの近似優先クエリキャッシュ(中間結果キャッシュ)。EH×スケッチで [[Prometheus]]/[[VictoriaMetrics]] のルールクエリを高速化・低コスト化(product / time-series)
- [[Prometheus]] — クラウドネイティブ時系列モニタリングの OSS de facto 標準。recording/alerting rule による周期ルールクエリを支援する TSDBMS(product / time-series)
- [[VictoriaMetrics]] — Prometheus 互換の高性能時系列モニタリングシステム。優れたストレージエンジン・並列クエリで高速だがクエリ冗長性は残す TSDBMS(product / time-series)
- [[Froot-NetSys promsketch]] — [[PromSketch]] の公式実装リポジトリ(github.com/Froot-NetSys/promsketch、Go 約 5K 行)(repository)
- [[Kexin Chu]] — [[eInfer]] 論文の筆頭著者([[University of Connecticut]])(person)
- [[Yiwei Yang]] — [[eGPU]] 論文の筆頭著者かつ [[eInfer]] 共著([[UC Santa Cruz]])。eBPF×GPU を横断する研究者(person)
- [[Shizhen Zhao]] — [[eInfer]] 論文の共著者([[Shanghai Jiao Tong University]])(person)
- [[Bohua Zou]] — [[ProfInfer]] 論文の筆頭著者([[Huawei Hilbert Research Center Dresden]]/[[TU Munich]])(person)
- [[Debayan Roy]] — [[ProfInfer]] 論文の責任著者([[Huawei Hilbert Research Center Dresden]])(person)
- [[Haibo Chen]] — [[ProfInfer]] と [[PICKER]] の共著者([[Huawei Central Software Institute]]/[[Shanghai Jiao Tong University]] IPADS)(person)
- [[Min Si]] — [[NCCLX]] 論文(Collective Communication for 100k+ GPUs)の筆頭著者([[Meta]])(person)
- [[Pavan Balaji]] — [[NCCLX]] 論文の共著者([[Meta]])(person)
- [[James Hongyi Zeng]] — [[NCCLX]] 論文の責任著者([[Meta]])(person)
- [[Sébastien Darche]] — GPU カーネル低オーバーヘッドトレース収集論文([[hip-analyzer]])の筆頭著者([[Polytechnique Montréal]] [[DORSAL lab]])(person)
- [[Michel R. Dagenais]] — 同 TOPC 論文の共著者・[[DORSAL lab]] 主宰([[Polytechnique Montréal]])(person)
- [[Mingcong Han]] — [[PICKER]] 論文の筆頭著者([[Shanghai Jiao Tong University]] IPADS)(person)
- [[Rong Chen]] — [[PICKER]] 論文の責任著者([[Shanghai Jiao Tong University]] IPADS)(person)
- [[Tong Yu]] — [[eGPU]] 論文の共著者([[Eunomia Inc]])(person)
- [[Andrew Quinn]] — [[eGPU]] 論文の共著者([[UC Santa Cruz]])(person)
- [[Hong Xu]] — [[Mycroft]] 論文の共著者(The Chinese University of Hong Kong)(person)
- [[UC Santa Cruz]] — [[eGPU]]・[[eInfer]] の著者が所属する米国の研究大学(organization)
- [[University of Connecticut]] — [[eInfer]] の著者 [[Kexin Chu]] らが所属する米国の研究大学(organization)
- [[University of Washington]] — [[eInfer]] の著者 Chenxingyu Zhao が所属する米国の研究大学(organization)
- [[Shanghai Jiao Tong University]] — [[eInfer]]・[[ProfInfer]]・[[PICKER]] に関与する中国の研究大学。IPADS を擁する(organization)
- [[Institute of Parallel and Distributed Systems]] — [[Shanghai Jiao Tong University]] 内の研究所(IPADS)。[[PICKER]] の著者が所属(organization)
- [[Meta]] — 10 万 GPU 超のクラスタを運用し集合通信ライブラリ [[NCCLX]]・[[CTran]] を開発する企業(organization)
- [[ByteDance Seed]] — ByteDance の研究組織。[[Mycroft]] の複数著者が所属(organization)
- [[Polytechnique Montréal]] — GPU カーネル低オーバーヘッドトレース収集論文の著者が所属するカナダの工科大学(organization)
- [[DORSAL lab]] — [[Polytechnique Montréal]] の研究室。[[Michel R. Dagenais]] が主宰し TOPC・[[hip-analyzer]] を手がける(organization)
- [[Huawei Hilbert Research Center Dresden]] — [[ProfInfer]] の著者が所属する Huawei のドイツ・ドレスデン研究拠点(organization)
- [[Huawei Central Software Institute]] — [[ProfInfer]] の著者 [[Haibo Chen]] が所属する Huawei の研究組織(organization)
- [[TU Munich]] — [[ProfInfer]] の著者が所属するドイツの工科大学(organization)
- [[Eunomia Inc]] — [[eGPU]] の著者 [[Tong Yu]] が所属し eGPU/bpftime のコードを公開する組織(organization)
- [[NCCLX]] — [[Meta]] の集合通信フレームワーク。[[NCCL]] を基盤に拡張し 10 万 GPU 超で訓練・推論を一元支援(product)
- [[CTran]] — [[NCCLX]] のカスタムトランスポート層。ゼロコピー・SM フリー・ホスト駆動(product)
- [[DQPLB]] — [[NCCLX]] の輻輳管理機構(Dynamic Queue Pair Load Balancing)(product)
- [[torchcomms]] — [[NCCLX]] の公開コードを含む [[Meta]] のリポジトリ(repository)
- [[Llama4]] — [[Meta]] の LLM。[[NCCLX]] の評価対象ワークロード(product)
- [[llama.cpp]] — エッジ/オンデバイス向け LLM 推論エンジン([[ProfInfer]] の計装対象)(repository)
- [[GGML]] — [[llama.cpp]] の ML ランタイム/テンソルライブラリ(repository)
- [[Perfetto]] — タイムライン可視化ツール。[[ProfInfer]] の ProfTime が利用(product)
- [[BCC]] — BPF Compiler Collection。eBPF 実装フレームワーク(repository)
- [[libbpf]] — eBPF 実装用 C ライブラリ(repository)
- [[Orange Pi]] — [[ProfInfer]] の評価デバイス(RK3588 系 SoC)(product)
- [[Rubik Pi]] — [[ProfInfer]] の評価デバイス(QCS6490 SoC)(product)
- [[Rockchip NPU]] — [[ProfInfer]] が自作バックエンドを実装した NPU(product)
- [[NVBit]] — NVIDIA GPU バイナリ計装フレームワーク。[[eGPU]] の比較対象(product)
- [[CUPTI]] — CUDA Profiling Tools Interface。GPU プロファイリングの比較参照(product)
- [[PTX]] — NVIDIA GPU の中間表現。[[eGPU]] の注入対象(product)
- [[ParallelGPU OS]] — GPU 向け並行チェックポイント/復元システム(POS)。[[eGPU]] のセルフモディファイングコード基盤(repository)
- [[CXLMemSim]] — CXL.mem シミュレータ([[eGPU]] 著者の先行研究)(repository)
- [[hip-analyzer]] — CUDA/[[HIP]] カーネル計装ツール(TOPC の参照実装)(repository)
- [[Rodinia]] — GPU ベンチマークスイート(TOPC の評価)(dataset)
- [[HIP]] — AMD の GPU プログラミングモデル(product)
- [[PICKER]] — インスタンス単位の GPU カーネル[[べき等性]]検証システム(product)
- [[Asymmetric Resilience]] — GPU 向け耐障害システム。[[PICKER]] の統合先(product)
- [[Chimera]] — プリエンプティブ GPU スケジューリングシステム。[[PICKER]] の統合先(product)
- [[Intel Corporation]] — 米国の半導体企業(Intel)。[[Aurora]] の自動故障管理 [[StabilityDB]] を [[RIKEN Center for Computational Science]] と共同開発(organization / hpc)
- [[RIKEN Center for Computational Science]] — 日本の計算科学研究拠点(R-CCS、神戸)。[[Intel Corporation]] と [[StabilityDB]] を共同研究(organization / hpc)
- [[Argonne National Laboratory]] — 米国エネルギー省の国立研究所(ANL)。エクサスケールスパコン [[Aurora]](63,744 GPU)を運用(organization / hpc)
- [[Store Foundational AI]] — [[Amazon Web Services]] の基盤 AI 部門。ストラグラー検知 + ノード健全性管理システム [[Guard]] を開発(organization / mlsys)
- [[iFLYTEK AI Engineering Institute]] — 中国・合肥の iFLYTEK の AI 工学研究所。[[FlashRecovery]] の主所属(organization / mlsys)
- [[Huawei Technologies]] — 中国の通信機器・半導体企業。[[Ascend NPU]] を擁し [[FlashRecovery]] を共同開発(organization)
- [[Beihang University]] — 中国・北京の研究大学(北京航空航天大学)。[[Hawkeye]] の共同研究機関(organization / networking)
- [[Infrawaves]] — [[Hawkeye]] に参加した企業([[Tsinghua University]]/[[Beihang University]] と共同、Xiaohe Hu の所属)(organization / networking)
- [[Hong Kong University of Science and Technology]] — 香港の研究大学(HKUST)。[[C4]] の共同研究機関([[Alibaba Group]] と共同)(organization)
- [[Nankai University]] — 中国・天津の研究大学(南開大学)。AIOps グループが光トランシーバー故障の予測 + 分類 [[OptProphet]] を開発(organization / aiops)
- [[Hewlett Packard Labs]] — HPE の研究部門。[[PACE]](ISAV、データセンター信頼性の因果探索)の主所属(organization)
- [[Oak Ridge National Laboratory]] — 米 DOE 国立研(ORNL)。Summit と冷却インフラの 7 年テレメトリを提供し [[PACE]] の評価データ側(organization / hpc)
- [[Case Western Reserve University]] — 米国の研究大学。GPU 性能モデリング論文の主所属(著者 5 名、case.edu)(organization / mlsys)
- [[Rutgers University]] — 米国の州立大学。GPU 性能モデリング論文の共同所属([[Mingkai Zheng]]・[[Zhao Zhang]]、rutgers.edu)(organization)
- [[Ant Group]] — 6,000 GPU 訓練クラスタで [[Flare]]([[XPUTimer]])を本番運用、[[DLRover]] の母体([[Ant International]] とは別)(organization / mlsys)
- [[Huawei Cloud]] — Computing and Networking Innovation Lab。本番 LLM 訓練基盤 [[Platform-X]] を運営し [[L4]]・[[LLMPrism]] の産業側所属(organization)
- [[Renmin University of China]] — 中国・北京の研究大学(中国人民大学)。[[StepFly]] 共著者 [[Yanjie Gao]] の所属([[Microsoft]] Research 兼任)(organization)
- [[Shuaiyu Xie]] — [[Wuhan University]] コンピュータ科学科。[[TVDiag]](マルチモーダル障害診断)の第一著者(person)
- [[Jian Wang]] — [[Wuhan University]] / Zhongguancun Laboratory。[[TVDiag]] 対応著者(person)
- [[Bing Li]] — [[Wuhan University]] / Zhongguancun Laboratory。[[TVDiag]] 対応著者(person)
- [[Wuhan University]] — 中国湖北省武漢市の総合研究大学(武漢大学)。[[TVDiag]] の主要所属機関(organization)
- [[TVDiag]] — [[Wuhan University]] の [[Shuaiyu Xie]]・[[Bing Li]] らが開発したマイクロサービス向けマルチモーダル障害診断フレームワーク(product)
- [[Xiao Zhang]] — [[Shandong University]] コンピュータ科学技術学部准教授。[[TAMO]] 第一著者。データマイニング・分散学習・エッジインテリジェンスを研究(person / aiops)
- [[Dongxiao Yu]] — [[Shandong University]] コンピュータ科学技術学部教授。[[TAMO]] 責任著者。エッジインテリジェンス・分散コンピューティング・データマイニングが専門(person / aiops)
- [[Fuzhen Zhuang]] — [[Beihang University]] 人工知能研究所教授。[[TAMO]] 共著者。転移学習・知識グラフ・推薦システムで 150 本超、CAAI 優秀論文賞 2013(person / ml)
- [[Shandong University]] — 中国・青島/済南の総合研究大学(山東大学、SDU)。[[Xiao Zhang]]・[[Dongxiao Yu]] らが所属し [[TAMO]] を開発(organization / aiops)
- [[DB-GPT]] — [[Tsinghua University]] の [[Guoliang Li]]・[[Xuanhe Zhou]] らが開発・公開する LLM 駆動データベース診断・最適化フレームワーク([[D-Bot]] の OSS 実装)(repository / database)
- [[Shiyue Huang]] — [[Peking University]] 博士課程学生。[[OpDiag]] 第一著者・[[DBPA]] 共同開発者。AI4DB・データベース異常診断(person)
- [[Bin Cui]] — [[Peking University]] 教授 IEEE Fellow。[[OpDiag]] 責任著者。データベース・機械学習・AI4DB(person)
- [[Yinjun Wu]] — [[Peking University]] 助教。Penn 博士(2021)・Tsinghua 学士(2016)。データ管理・機械学習(person)
- [[Ziwei Wang]] — [[Hong Kong University of Science and Technology, Guangzhou]] 博士課程学生(PKU 修士 2024)。[[OpDiag]] 共著者(person)
- [[ZTE Corporation]] — 中国・南京の通信 ICT 企業。[[OpDiag]] 産業パートナー・[[DBPA]] DBA 提供(organization)
- [[OpDiag]] — [[Peking University]]/[[ZTE Corporation]] のクエリ演算子レベル DB 性能異常診断フレームワーク(product)
- [[DBPA]] — PKU/ZTE による OLTP DB 性能異常再現ベンチマーク(ACM SIGMOD 2023)(dataset)
- [[AgentTune]] — Renmin University of China/ByteDance の LLM エージェントベースデータベースノブチューニングフレームワーク(4 エージェント + ビームサーチ木探索、全実験で Invalid Times=0、SIGMOD 2025)(product / database)
- [[Yiyan Li]] — AgentTune 第一著者(Renmin University of China、データベースノブチューニング×LLM)(person / database)
- [[Haoyang Li]] — AgentTune 第一著者(Renmin University of China、ノブチューニング・NL2SQL)(person / database)
- [[Jing Zhang]] — AgentTune 共著者(Renmin University of China)(person / database)
- [[Cuiping Li]] — AgentTune 責任著者(Renmin University of China 教授、MOE キーラボ・研究センター)(person / database)
- [[Hong Chen]] — AgentTune 共著者(Renmin University of China 教授、MOE キーラボ)(person / database)
- [[Renata Borovica-Gajic]] — AgentTune 共著者(University of Melbourne、データベースシステム・チューニング、ARC DECRA DE230100366)(person / database)
- [[University of Melbourne]] — オーストラリア・メルボルンの研究大学。AgentTune 共著者 Renata Borovica-Gajic の所属(organization)
- [[SCELM]] — ECD・FT・RCCA 統合ソフトウェア変更評価フレームワーク。Nankai University AIOps グループ開発(product)
- [[Tinghua Zheng]] — Nankai University 所属の研究者。SCELM 共著者(person)
- [[Xidao Wen]] — BizSeer 所属の研究者。SCELM 共著者(person)
- [[Weihua Kuang]] — Nankai University 所属の研究者。SCELM 共著者(person)
- [[Heng Liu]] — CHINA TIANCHEN ENGINEERING CORPORATION LTD. 所属の研究者。SCELM 共著者(person)
- [[Chao Shen]] — Nankai University 所属の研究者。SCELM 共著者(person)
- [[Bo Wu]] — Tencent Technologies 所属の研究者。SCELM 共著者(person)
- [[BizSeer]] — 北京の企業。AIOps 研究開発。Xidao Wen の所属(organization)
- [[Binpeng Shi]] — Nankai University College of Software の研究者。FlowXpert 第一著者(person / aiops)
- [[FlowXpert]] — トラブルシューティングワークフロー自動生成フレームワーク。Nankai University + Huawei Cloud 共同開発、10 週間本番展開・承認率約 80%(product / incident-management)
- [[Kaleidoscope]] — UIUC/NCSA の Jha ら(SC 2020)が開発した HPC 分散ストレージ向け近リアルタイム障害フォレンジクスフレームワーク。Blue Waters に実導入(product / hpc / aiops)
- [[Blue Waters]] — NCSA/UIUC が運用するペタスケール HPC スーパーコンピュータ。Cray Sonexion(Lustre、36 PB)付き。Kaleidoscope の評価環境(product / hpc)
- [[Subho S. Banerjee]] — UIUC の Kaleidoscope 共著者(person)
- [[Zbigniew T. Kalbarczyk]] — UIUC のディペンダブルコンピューティング研究者。Kaleidoscope 共著(person / hpc)
- [[OpsFlowBench]] — Huawei Cloud データセンタースイッチ運用ドキュメント由来の 252 件ワークフロー評価ベンチマーク(dataset / aiops)
- [[Haoming Meng]] — CUJBench の単著著者。所属記載なし(person)
- [[CUJBench]] — ブラウザ可視証拠+バックエンド可観測性の初のクロスモーダル障害診断ベンチマーク。87 シナリオ・A@1=19.7%・天井=52%(product / benchmark)
- [[OpenTelemetry Demo]] — OpenTelemetry プロジェクトのポリグロットマイクロサービス型 EC デモ。CUJBench のテスト環境の 1 つ(product)
- [[Tractor Store]] — マイクロフロントエンド型オープンソース EC アプリ。CUJBench のテスト環境の 1 つ(product)
- [[Lustre]] — オープンソース並列ファイルシステム(GPL-2.0, 750K LOC)。Top500 上位 10 中 6 台・上位 500 の 60% 超が採用。25 年超の開発史(product / hpc / storage)
- [[Frontier]] — ORNL のエクサスケールスパコン(AMD EPYC + MI250X、Top500 2024-06 #1)。Lustre FS [[Orion]] を運用(product / hpc)
- [[Orion]] — Frontier の Lustre ファイルシステム。700 PB・フラッシュ+HDD 多層・40 MDT、逐次読み出し 4.7 TiB/s(product / hpc / storage)
- [[DDN]] — HPC ストレージベンダ(旧 DataDirect Networks)。2018 年に Intel の Lustre 事業を買収し [[Whamcloud]] を傘下に(organization / hpc)
- [[Whamcloud]] — Lustre 開発企業。元 CFS 開発者が設立、Intel(2012)→DDN(2018)傘下(organization / hpc)
- [[OpenSFS]] — Open Scalable File Systems, Inc.。Lustre オープンソースコミュニティの推進コンソーシアム(organization / hpc)
- [[Anjus George]] — ORNL NCCS 所属。Lustre Unveiled(TOS 2025)筆頭著者(person / hpc)
- [[Andreas Dilger]] — Whamcloud/DDN 所属。Lustre の主要コントリビュータ(person / hpc)
- [[Sarp Oral]] — ORNL NCCS 所属。Frontier/Orion ストレージリーダー(person / hpc)
- [[Peter J. Braam]] — Lustre の創出者。CMU Coda プロジェクト出身、[[Cluster File Systems]] を設立(person / hpc)
- [[Cluster File Systems]] — Lustre を開発したスタートアップ(CFS)。Sun Microsystems(2007)→Oracle(2010)に買収(organization / hpc)
- [[Yu Guan]] — EROICA 筆頭著者(Alibaba Cloud / Zhejiang Lab)(person / mlsys)
- [[Dan Li]] — Wormhole(SimNet)共著者(Tsinghua University、ネットワークシミュレーション)(person / networking)
- [[Fuheng Zhao]] — PrvTel 共著者(University of Maryland、差分プライバシー×テレメトリ)(person)
- [[Zhejiang Lab]] — 中国・杭州の研究機関(之江実験室)。EROICA の共同研究機関(organization / mlsys)
- [[Zhongguancun Laboratory]] — 中国・北京の研究機関(中関村実験室)。Wormhole(SimNet)の共同研究機関(organization / networking)
- [[MangoBoost]] — 韓国のスタートアップ。FAST スケジューラの共同研究機関(organization / networking)
- [[University of Pennsylvania]] — 米国の研究大学(UPenn)。FAST 共著者の所属(organization)
- [[Northeastern University]] — 中国・遼寧省瀋陽市の重点理工系大学。HeteCCL の主要所属機関(organization)
- [[Shenzhen Institutes of Advanced Technology]] — 中国科学院傘下の深圳市研究機関(SIAT)。HeteCCL の共同研究機関(organization)
- [[Max Planck Institute for Informatics]] — ドイツ・ザールブリュッケンの Max-Planck 傘下情報科学研究所。Matryoshka 共著(organization)
- [[Max Planck Institute for Software Systems]] — ドイツ・ザールブリュッケンの Max-Planck 傘下ソフトウェアシステム研究所。[[Hindsight]] の著者所属(organization)
- [[Agent-R1]] — USTC の [[Mingyue Cheng]] らが開発したエージェント型 RL 訓練フレームワーク。ステップレベル MDP 抽象化と柔軟なコンテキスト管理を核に PPO・GRPO・Reinforce++・RLOO を比較(product / agentic-rl)
- [[AutoForge]] — Tongyi Lab(Alibaba Group)のエージェント型 RL フレームワーク。ツール記述文書から模擬環境を自動合成し ERPO で訓練(product / agentic-rl)
- [[Tongyi Lab]] — Alibaba Group の AI 研究組織。Qwen シリーズ・AutoForge を開発(organization)
- [[Fuli Feng]] — Tongyi Lab(Alibaba Group)の研究者。AutoForge 共著者(person / agentic-rl)
- [[DeepSWE]] — Agentica/Together AI の完全オープンソース RL 訓練コーディングエージェント。Qwen3-32B + GRPO++ で SWE-Bench-Verified Pass@1 42.2%・ハイブリッド Best@16 59.0%(product / coding-agent)
- [[Together AI]] — AI インフラストラクチャ企業。DeepSWE の共同開発元(organization)
- [[Agentica]] — 言語エージェント事後学習フレームワーク rLLM を開発する研究チーム。DeepSWE の共同開発元(organization)
- [[Ion Stoica]] — UC Berkeley 教授。Spark/Ray/Anyscale 共同創設者、DeepSWE 共著者(person)
- [[Raluca Ada Popa]] — UC Berkeley 教授。セキュリティ・ML システム、DeepSWE 共著者(person)
- [[rLLM]] — Agentica の言語エージェント事後学習フレームワーク。GRPO++ と Kubernetes 統合を実装、DeepSWE の訓練基盤(product / agentic-rl)
- [[R2E-Gym]] — ソフトウェアエンジニアリング問題の RL 訓練用データセット/環境。DeepSWE で 4,500 問サブセットを使用(dataset)
- [[SWE-Bench-Verified]] — ソフトウェアエンジニアリング自動化の標準ベンチマーク。DeepSWE がオープンウェイト SOTA を達成(dataset)
- [[Michael Luo]] — DeepSWE 筆頭著者。Agentica でモデル訓練・インフラを主導(person / coding-agent)
- [[Naman Jain]] — DeepSWE 共同筆頭著者。R2E-Gym 開発・エージェントスキャフォールド設計を担当(person / coding-agent)
- [[Aviral Kumar]] — Carnegie Mellon University の研究者。IsoCompute Playbook の corresponding author。オフライン RL(CQL)・RL スケーリング則の研究者(person / reinforcement-learning)
- [[Zhiting Hu]] — UC San Diego 准教授。IsoCompute Playbook のシニア著者。テキスト生成・構造化 LLM・制御可能生成(person / machine-learning)
- [[MBZUAI]] — Mohamed bin Zayed University of Artificial Intelligence。アブダビの AI 専門大学院大学。IsoCompute Playbook の共著機関(MBZUAI-IFM)(organization)
- [[Zelin Tan]] — RL スケーリング則論文の筆頭著者(USTC / Shanghai AI Lab インターン)(person / reinforcement-learning)
- [[Chen Zhang (Shanghai AI Lab)]] — RL スケーリング則論文の責任著者(Shanghai AI Laboratory)。同名別人物あり(person / reinforcement-learning)
- [[Zhenfei Yin]] — RL スケーリング則論文の責任著者(University of Oxford、Philip Torr グループ)(person / reinforcement-learning)
- [[University of Oxford]] — 英国の研究大学。RL スケーリング則研究の共同研究機関(organization)
- [[VeRL]] — LLM 向け大規模 RL 訓練プラットフォーム(HybridFlow, EuroSys 2025)。RL スケーリング則の全実験基盤(product / reinforcement-learning)
- [[GRPO]] — Group Relative Policy Optimization。グループ内報酬正規化でアドバンテージを推定するアクターのみの RL アルゴリズム(DeepSeekMath, Shao+ 2024)。RL スケーリング則・DeepSWE・AutoForge 等で広く使用(product / reinforcement-learning)
- [[AgentRL]] — [[Tsinghua University]]/[[Z.AI]] のマルチターン・マルチタスク・エージェント型 RL 訓練フレームワーク。交差方策サンプリング・タスク別アドバンテージ正規化・完全非同期パイプラインで 5 環境平均成功率 70.4%(product / agentic-rl)
- [[Hanchen Zhang]] — [[Tsinghua University]]/[[Z.AI]] の研究者。AgentRL 共同筆頭著者(person / agentic-rl)
- [[Xiao Liu]] — [[Tsinghua University]]/[[Z.AI]] の研究者。AgentRL 共同筆頭著者。AgentBench・AutoGLM の主導者(person / agentic-rl)
- [[Yuxiao Dong]] — [[Tsinghua University]] 准教授。AgentRL 責任著者。ウェブマイニング・ソーシャルネットワーク分析(person / agentic-rl)
- [[Z.AI]] — AgentRL 共同研究機関。[[Tsinghua University]] と連携し AgentRL・AutoGLM を開発(organization)
- [[AgentBench]] — [[Xiao Liu]] ら([[Tsinghua University]])が開発した LLM エージェントのマルチターン評価ベンチマーク(ICLR 2024)。AgentRL で AgentBench-FC として使用(dataset)
- [[Devvrit Khatri]] — UT Austin PhD(Inderjit S. Dhillon グループ)。ScaleRL 論文の共同筆頭著者(person / reinforcement-learning)
- [[Rishabh Agarwal]] — Periodic Labs の研究者。ScaleRL 論文の責任著者。深層 RL の統計的評価(rliable)で著名(person / reinforcement-learning)
- [[ScaleRL]] — Meta/UT Austin/UC Berkeley/Harvard/Periodic Labs の RL 訓練レシピ。PipelineRL-8 + CISPO + FP32 精度修正 + 損失集約 + アドバンテージ正規化 + ゼロ分散フィルタリング + No-Positive-Resampling の 7 成分で 8B A=0.61・Scout 17B×16 MoE A=0.71(product / reinforcement-learning)
- [[PipelineRL]] — Piche ほか(2025)の LLM RL 訓練向けストリーミング型非同期パイプラインセットアップ。ScaleRL の基盤コンポーネント(product / reinforcement-learning)
- [[UT Austin]] — University of Texas at Austin。米国テキサス州オースティンの州立研究大学(organization)
- [[Mingjie Liu]] — NVIDIA の研究者。Scaling Up RL 論文の筆頭著者。LLM の長期 RL 訓練を主導(person / reinforcement-learning)
- [[Yejin Choi]] — NVIDIA / University of Washington 教授。Scaling Up RL 論文の著名共著者。NLP・常識推論の著名研究者(person / reinforcement-learning)
- [[NVIDIA]] — GPU・AI アクセラレータ・AI インフラの最大手企業。H100/H200 等の GPU、NCCL 等の集団通信ライブラリ、Megatron-LM 等の訓練フレームワーク、TensorRT-LLM/NVIDIA NIM 等の推論スタック、NIXL 等のデータ転送ライブラリを開発。Scaling Up RL で Nemotron-Research-Reasoning-Qwen-1.5B を公開(organization)
- [[Nemotron-Research-Reasoning-Qwen-1.5B]] — NVIDIA が DeepSeek-R1-Distill-Qwen-1.5B をベースに 5 ドメイン長期 RL で訓練した 1.5B パラメータ推論特化モデル。数学 +14.7%・コード +13.9%・論理パズル +54.8%(product / reinforcement-learning)
- [[Lovish Madaan]] — Meta/UCL の研究者。ScaleRL 論文の共同筆頭著者・対応著者(person / reinforcement-learning)
- [[David Brandfonbrener]] — Meta の研究者。ScaleRL 論文の対応著者(person / reinforcement-learning)
- [[Periodic Labs]] — Rishabh Agarwal の所属組織。ScaleRL 論文の責任著者機関(organization)
- [[Philip Torr]] — University of Oxford 教授。コンピュータビジョン・機械学習。RL スケーリング則論文の共著者(person / reinforcement-learning)
- [[Lei Bai]] — Shanghai AI Laboratory の研究者。RL スケーリング則論文・Landscape サーベイの責任著者(person / reinforcement-learning)
- [[Guibin Zhang]] — National University of Singapore の研究者。Landscape of Agentic RL サーベイの共同筆頭著者(person / agentic-rl)
- [[Zhoujun Cheng]] — UCSD/MBZUAI-IFM の研究者。IsoCompute Playbook の共同筆頭著者(person / reinforcement-learning)
- [[Alexander Golubev]] — Nebius AI の研究者。Training SWE Agents 論文の筆頭著者・対応著者(person / coding-agent)
- [[Nebius AI]] — 旧 Yandex Cloud。SWE エージェントの RL 訓練研究の主所属機関。H200 クラスタ運用(organization)
- [[Boris Yangel]] — Humanoid/Nebius AI の研究者。Training SWE Agents 論文の最終著者(person / coding-agent)
- [[Humanoid]] — Boris Yangel の投稿時所属組織(organization)
- [[Ce Zhang]] — Together AI 共同創設者。元 ETH Zurich 教授。DeepSWE 共著者(person / machine-learning)
- [[Koushik Sen]] — UC Berkeley 教授。プログラム解析の著名研究者。DeepSWE 共著者(person / software-engineering)
- [[Li Erran Li]] — Together AI の研究者。元 AWS AI Labs。DeepSWE 共著者(person / machine-learning)
- [[Archit Patke]] — Delta A100 GPU レジリエンス研究の共同筆頭著者(UIUC)(person / hpc)
- [[Ziheng Chen]] — Delta A100 GPU レジリエンス研究の共同筆頭著者(UIUC)(person / hpc)
- [[Aditya Ranjan]] — Delta A100 GPU レジリエンス研究の共同筆頭著者(UIUC)(person / hpc)
- [[Hung Nguyen]] — Delta A100 GPU レジリエンス研究の共著者(UIUC)(person / hpc)
- [[Phuong Cao]] — Delta A100 GPU レジリエンス研究の共著者(UIUC)(person / hpc)
- [[Brett Bode]] — Delta A100 GPU レジリエンス研究の共著者(UIUC/NCSA)(person / hpc)
- [[Gregory Bauer]] — Delta A100 GPU レジリエンス研究の共著者(UIUC/NCSA)(person / hpc)
- [[Chandra Narayanaswami]] — Delta A100 GPU レジリエンス研究の共著者(IBM Research)(person / hpc)
- [[Daby Sow]] — Delta A100 GPU レジリエンス研究の共著者(IBM Research)(person / hpc)
- [[Catello Di Martino]] — Delta A100 GPU レジリエンス研究の共著者(Nokia Bell Labs)(person / hpc)
- [[Cursor]] — AI コーディングエディタ / エージェントモデルプロバイダ。Composer シリーズを提供し、Moonshot の Kimi K2.5 を基盤にターゲット RL で訓練(product / coding-agent)
- [[Kimi K2.5]] — Moonshot のオープンソース LLM チェックポイント。Cursor Composer 2/2.5 の基盤モデル(product / foundation-model)
- [[Moonshot AI]] — 中国の AI 企業(月之暗面)。Kimi シリーズ LLM を開発。Kimi K1.5/K2/K2.5 および Kimi-Researcher を公開(organization)
- [[SpaceXAI]] — Cursor と協業し大規模モデル訓練を推進。Colossus 2 の百万 H100 相当の計算資源を保有(organization)
- [[Colossus 2]] — 百万 H100 相当の GPU を擁する大規模 AI 訓練クラスタ。SpaceXAI が運用(product / hpc)
- [[Sharded Muon]] — 分散直交化と dual-mesh HSDP を統合した LLM 訓練オプティマイザ。Cursor Composer 2.5 の訓練スタックに採用(product / optimization)
- [[Kimi K2]] — [[Moonshot AI]] の 1.04 兆パラメータ(活性化 32B)超疎 MoE LLM。384 エキスパート・MLA・MuonClip で 15.5 兆トークン事前学習。SWE-bench Verified 65.8%(product / foundation-model)
- [[MuonClip]] — [[Moonshot AI]] が [[Kimi K2]] 訓練用に開発した Muon + QK-Clip オプティマイザ。ヘッド単位のロジットクリッピングで 1T MoE の事前学習を安定化(product / optimization)
- [[NorMuon]] — [[Datadog]] AI Research が Toto 2.0 訓練用に開発した per-neuron 正規化 Muon オプティマイザ。ピンボール損失の符号値勾配問題を EMA スケール row normalization で解決(product / optimization)
## Organization
- [[MiniMax]] — 中国の AI 企業。ハイブリッドアテンション推論モデル [[MiniMax-M1]] および MoE 基盤の [[MiniMax-M2]] シリーズの開発元(organization)
## Product / System
- [[MiniMax-M1]] — [[MiniMax]] の 456B/45.9B ハイブリッドアテンション推論モデル。ライトニングアテンション 7:ソフトマックス 1 の混成で 100 万トークンコンテキスト・100K 思考予算を実現(product)
- [[MiniMax-Text-01]] — [[MiniMax-M1]] のベースモデル。456B 総パラメータ・45.9B 活性化・32 エキスパートの MoE ハイブリッドアテンションアーキテクチャ(product)
- [[CISPO]] — Clipped IS-weight Policy Optimization。重要度サンプリング重みをクリッピングする RL アルゴリズム。ハイブリッドアテンションモデルの精度不一致に起因する省察トークンクリッピング問題を解決し DAPO 比 2 倍のステップ効率(product)
- [[Lightning Attention]] — I/O 認識型の線形アテンション実装(Qin+ 2024b)。ソフトマックスアテンションの二次計算量をバイパスし長系列生成の FLOPS を大幅削減(product)
- [[SynLogic]] — [[MiniMax]] の論理推論データ合成フレームワーク。41 タスクファミリーを自動生成し RL 訓練の初期データとして使用(product)
- [[MiniMax-M2]] — [[MiniMax]] の 229.9B/9.8B MoE 言語モデルファミリー。256 細粒度エキスパート・シグモイドゲーティング・192K コンテキスト・MTP 投機的復号。M2→M2.5→M2.7 の自己進化(product)
- [[Forge]] — [[MiniMax]] のエージェントネイティブ RL 訓練インフラ。Agent/Middleware/Training の 3 モジュール分離、ホワイトボックス/ブラックボックスエージェント統一、Windowed FIFO、接頭辞木マージで最大 40× 高速化(product)
- [[Nemotron 3]] — [[NVIDIA]] のオープン LLM ファミリー。ハイブリッド Mamba-2–Transformer MoE アーキテクチャ、[[LatentMoE]]、NVFP4 事前学習、マルチ環境同時 RL を特徴とする。Nano(30B/3B)・Super・Ultra の 3 モデル(product)
- [[LatentMoE]] — [[NVIDIA]] が [[Nemotron 3]] で提案したハードウェア認識型 MoE エキスパート設計。潜在次元でのエキスパート計算と通信削減を精度向上に再投資(product)
- [[NeMo-RL]] — [[NVIDIA]] のスケーラブル RL 訓練フレームワーク(Apache 2.0)。Nemotron 3 のマルチ環境 RL ポストトレーニングの基盤(product)
- [[Kimi K1.5]] — [[Moonshot AI]] の RL 訓練マルチモーダル LLM。長コンテキスト RL(128k)とオンラインミラー降下変種で OpenAI o1 に匹敵する推論性能を達成。パーシャルロールアウトと long2short 手法を導入(product / foundation-model)
- [[Mooncake]] — [[Moonshot AI]] の RDMA ベースチェックポイント転送エンジン。[[Kimi K1.5]] の RL インフラで [[Megatron-LM]] から [[vLLM]] への重み転送に使用(product / mlsys)
- [[Kimi-Researcher]] — [[Moonshot AI]] のエンドツーエンド RL 訓練による自律型リサーチエージェント。REINFORCE のみで HLE Pass@1 26.9%・xbench-DeepSearch 69%(product / agent)
- [[Cursor Research]] — AI コーディングエディタ Cursor のリサーチ部門。Composer シリーズ(Composer 2/2.5)を開発(organization)
- [[Fireworks AI]] — LLM 推論基盤企業。Composer 2 の RL 推論パートナーとして地理的に分散した推論クラスタを運用(organization)
- [[Composer 2]] — [[Cursor Research]] のエージェント型コーディングモデル(1.04T/32B MoE、Kimi K2.5 ベース)。CursorBench 61.3・SWE-bench Multi 73.7 でコスト精度パレート最適(product)
- [[CursorBench]] — Cursor の内部コーディングエージェントベンチマーク。中央値 181 行変更の実世界 Cursor セッション由来(dataset)
- [[Anyrun]] — Cursor の Firecracker VM 基盤コード実行プラットフォーム。RL 環境兼クラウドエージェント実行基盤(product)
- [[ThunderKittens]] — Stanford Hazy Research の GPU カーネルフレームワーク。BF16/MXFP8/NVFP4 GEMM・Flash Attention 4 を提供(repository)
- [[DeepEP]] — DeepSeek のエキスパート並列通信ライブラリ。MoE のトークンディスパッチ/コンバインを効率化(repository)
## Organization
- [[Allen Institute for AI]] — 米国シアトルの非営利 AI 研究所(AI2)。完全オープン LLM [[OLMo 3]] を開発し、モデルフロー全体(全段階の重み・データ・コード・ログ)を公開(organization / open-lm)
## Product / System
- [[OLMo 3]] — [[Allen Institute for AI]] の完全オープン LLM ファミリー(7B/32B)。SWA アーキテクチャ、Base・Think・Instruct・RL-Zero の 4 変種。OLMo 3.1 Think 32B は MATH 96.2・AIME 2024 80.6 で完全オープンモデル最強(product / open-lm)
- [[OlmoRL]] — [[Allen Institute for AI]] の LLM 向け完全非同期 RL 訓練インフラ。[[GRPO]] ベース 7 改善、[[DeepSpeed]] + [[vLLM]] 非同期パイプライン、OLMo 2 比 4 倍スループット(product / reinforcement-learning)
- [[OlmoBaseEval]] — [[Allen Institute for AI]] のベースモデル評価スイート。43 タスク・5 クラスタ、タスククラスタリング・スケーリング分析・SNR 分析の 3 手法(product / evaluation)
- [[olmOCR]] — [[Allen Institute for AI]] の PDF テキスト変換パイプライン。2.38 億件のユニーク PDF を処理し [[Dolma 3]] の学術コーパスを構築(product / data-processing)
- [[Duplodocus]] — [[Allen Institute for AI]] の Rust 製大規模重複排除ツールキット。完全一致→MinHash→サフィックス配列の 3 段階で兆トークン規模のコーパスを処理(product / data-processing)
## Repository / Dataset
- [[Dolma 3]] — [[Allen Institute for AI]] の [[OLMo 3]] 事前学習データスイート。Dolma 3 Mix(5.9T)・Dolmino Mix(100B)・Longmino Mix(50–100B)の 3 段階構成(dataset / open-lm)
- [[Dolci]] — [[Allen Institute for AI]] の [[OLMo 3]] 後訓練データスイート。Think・Instruct・RL-Zero の 3 変種に SFT・DPO・RL データを提供。Delta Learning 用選好ペアを含む(dataset / post-training)
## Person
- [[Jim Gray]] — Tandem Computers の研究者。耐障害システムの古典論文 "Why Do Computers Stop" (1985) の著者。Bohrbug/Heisenbug 概念、プロセスペア類型を体系化。チューリング賞(1998)受賞者
- [[David Oppenheimer]] — UC Berkeley ROC Project の研究者。インターネットサービス障害分析論文 (USITS 2003) の筆頭著者
- [[Archana Ganapathi]] — UC Berkeley ROC Project の研究者。インターネットサービス障害分析論文 (USITS 2003) の共著者
- [[David A. Patterson]] — UC Berkeley 教授。RISC アーキテクチャ・RAID の共同発明者。ROC Project 主宰。インターネットサービス障害分析論文の共著者
- [[James Hamilton]] — Microsoft Windows Live Services Platform。インターネットスケールサービス設計のベストプラクティス論文 (LISA 2007) の著者。後に Amazon VP/Distinguished Engineer
- [[Lisanne Bainbridge]] — University College London 心理学部。"Ironies of Automation" (1983) の著者。自動化のパラドクスを人間工学の観点から体系化
- [[Ben Treynor Sloss]] — [[Google]] VP Engineering、SRE 創設者。「信頼性はあらゆるプロダクトの最も基本的な特性」
- [[Betsy Beyer]] — [[Google]] テクニカルライター。[[SRE Book]](2016)および The Site Reliability Workbook(2018)の編者
- [[Niall Murphy]] — [[Google]] SRE。[[SRE Book]] 共同編者、第 4 章(SLO)・第 7 章(自動化)の著者
- [[Chris Jones]] — [[Google]] App Engine SRE。[[SRE Book]] 第 4 章(SLO)共著者。SREcon16 でエラーバジェット制御ループを口頭解説
- [[Margaret Hamilton]] — MIT Instrumentation Laboratory ソフトウェアエンジニアリング部門ディレクター。Apollo 計画、「ソフトウェアエンジニアリング」の造語者
- [[Shutian Luo]] — SIAT/CAS 研究者。Alibaba トレース分析によるマイクロサービス依存関係・性能特性分析(SoCC 2021)
- [[Chengzhong Xu]] — University of Macau 教授。分散システム・クラウドコンピューティング。SoCC 2021 マイクロサービス特性分析の共著者
- [[Supriyo Ghosh]] — Microsoft 所属研究者。SoCC 2022 インシデント実証研究の筆頭著者
- [[Chetan Bansal]] — Microsoft 研究者。FLASH・SoCC 2022 インシデント研究の共著者。クラウドインシデント管理の AI 自動化
- [[Suman Nath]] — Microsoft 所属研究者。SoCC 2022 インシデント実証研究の共著者
- [[I-Ting Angelina Lee]] — Washington University in St. Louis 教員。Uber の非致命的 RPC エラー分析(PACM CAS 2024)
- [[Milind Chabbi]] — Uber Technologies リサーチサイエンティスト。CRISP・非致命的 RPC エラー研究の中心著者
- [[Darby Huye]] — Tufts University 博士課程(Meta 兼任)。Meta マイクロサービスアーキテクチャ分析(ATC 2023)
- [[Yuri Shkuro]] — Meta エンジニア。Jaeger(CNCF 分散トレーシング)の作者。Meta マイクロサービス分析の共著者
- [[Raja R. Sambasivan]] — Tufts University 准教授。分散システムの性能デバッグ。Meta マイクロサービス分析の共著者
- [[Zhe Xie]] — Tsinghua University / BNRist 研究者。LatentScope(KDD 2024)の筆頭著者
- [[Junxian Shen]] — Tsinghua University。DeepFlow(SIGCOMM 2023)の筆頭著者
- [[Han Zhang (DeepFlow)]] — Tsinghua University。DeepFlow(SIGCOMM 2023)の責任著者
- [[Nengwen Zhao]] — Tsinghua University & BNRist 研究者。SCWarn の筆頭著者(ESEC/FSE 2021)
- [[Zhizhou Zhang]] — UC Santa Barbara。CRISP(USENIX ATC 2022)と非致命的 RPC エラー研究(PACM CAS 2024)の筆頭著者
- [[Arvind Krishnamurthy]] — University of Washington 教授 / Google。分散システム・ネットワーキング。Google RPC 特性分析(SOSP 2023)の共著者
## Organization
- [[Tandem Computers]] — 耐障害コンピュータの先駆企業。NonStop システムを開発。1997 年 Compaq に買収
- [[UC Berkeley ROC Project]] — UC Berkeley の Recovery-Oriented Computing プロジェクト。David Patterson・Armando Fox が主導
- [[University College London]] — ロンドンの研究大学(UCL)。Lisanne Bainbridge の所属
- [[University of Macau]] — マカオの研究大学。分散システム・クラウドコンピューティング研究
- [[SIAT]] — 中国科学院傘下の研究機関(深圳)。分散システム・マイクロサービス研究
- [[Tufts University]] — 米国マサチューセッツ州の私立研究大学。Raja R. Sambasivan・Darby Huye の所属
- [[Washington University in St. Louis]] — 米国ミズーリ州の研究大学。I-Ting Angelina Lee の所属
- [[Yunshan Networks]] — DeepFlow の共同開発元。清華大学との産学連携
- [[China Guangfa Bank]] — 中国の大手商業銀行。SCWarn 研究の実データ提供元
- [[Alibaba Group]] — 企業(AIOps/LLM 研究、Qwen シリーズ開発元、マイクロサービストレース研究のデータ提供元)
- [[Uber]] — 非致命的 RPC エラー研究・CRISP の対象企業。6,000 超のマイクロサービスを運用
## Product / System
- [[NonStop]] — Tandem Computers の耐障害コンピュータシステム。プロセスペア・フェイルファスト・トランザクション機構による高可用性
- [[SRE Book]] — "Site Reliability Engineering: How Google Runs Production Systems"(O'Reilly, 2016)。SRE ディシプリンの定義書。全 34 章・3 部構成
- [[Canopy]] — Meta の分散トレーシング基盤(エンドツーエンド)。マイクロサービストポロジ分析の基盤
- [[DeepFlow]] — eBPF ベースのネットワーク中心分散トレーシングシステム。コード修正ゼロ・暗黙のコンテキスト伝搬
- [[SCWarn]] — 異種多ソースデータを統合して不正ソフトウェア変更を検出するマルチモーダル異常検知システム
- [[CRISP]] — マイクロサービス RPC トレースのクリティカルパス分析システム。Uber に実投入
- [[LatentScope]] — 限定観測可能性下のマイクロサービス RCA 手法。介入認識による潜在空間上の因果推論
- [[Microsoft Teams]] — Microsoft のリアルタイムコミュニケーションサービス。インシデント研究の対象サービス
- [[Dapper]] — Google の内部分散トレーシングサービス。RPC 特性分析の基盤
- [[Stubby]] — Google 内部の RPC ライブラリ。gRPC と同等の機能
- [[Monarch]] — [[Google]] が 2010 年から運用するプラネットスケール・マルチテナント・インメモリ時系列データベース。2019 年 7 月時点で約 950 億時系列・750 TB インメモリ・2.2 TB/s 取り込み。リレーショナルデータモデル・FHI(トライグラムインデックス)・Collection Aggregation・クエリプッシュダウンを主要技術とする。前身は [[Borgmon]] (product / time-series / observability)
- [[Borgmon]] — [[Google]] の初代分散監視システム。[[Monarch]] の前身。pull 型・push 型混在で各チームが独立運用するデーモン型アーキテクチャ。スキーマなし・distribution 型非サポート・手動シャーディングの 4 課題が Monarch 設計の動機となった。Prometheus 等 OSS 監視ツールの設計に影響を与えた (product / observability)
- [[Jaeger]] — CNCF の分散トレーシングプラットフォーム。Yuri Shkuro が開発
## DeepSeek ファミリー
### Person
- [[Daya Guo]] — DeepSeek-Coder 筆頭著者。コード LLM 研究の中心人物
### Organization
- [[DeepSeek-AI]] — 中国の AI 研究企業。DeepSeek LLM/Coder/V3/R1/V3.2/VL2/V4 シリーズを開発。MoE アーキテクチャと効率的訓練で知られる
### Product / System
- [[DeepSeek LLM]] — DeepSeek 初代基盤モデル(7B/67B dense)。非埋め込み FLOPS/トークン M によるスケーリング則研究
- [[DeepSeek-Coder]] — DeepSeek のコード特化 LLM(1.3B〜33B)。87 言語・2T トークン、FIM 最適化。6.7B で CodeLlama-34B を凌駕
- [[DeepSeek-V3]] — DeepSeek の 671B MoE モデル(37B 活性化)。MLA・補助損失なし負荷分散・MTP・FP8・DualPipe
- [[DeepSeek-R1]] — DeepSeek の推論特化モデル。SFT なし純粋 RL で推論能力の創発を実証
- [[DeepSeek-R1-Zero]] — DeepSeek-R1 の SFT なし純粋 RL 変種。aha モーメントの創発で注目
- [[DeepSeek-V3.2]] — V3 後継。DSA・GRPO 安定化・1,800+ 合成エージェント環境で事後学習を大幅強化
- [[DeepSeek-VL2]] — DeepSeek の MoE ベース VLM(3B/4.5B/27B 活性化)。動的タイリング視覚エンコーダ
- [[DeepSeek-V4]] — DeepSeek 最新フラグシップ。MegaMoE・CSA+HCA ハイブリッド圧縮アテンションで 100 万トークンコンテキスト
- [[Multi-head Latent Attention]] — DeepSeek-V3 が導入した低ランク KV 圧縮アテンション機構(MLA)
- [[DualPipe]] — DeepSeek-V3 のパイプライン並列手法。計算と通信の双方向オーバーラップでバブル率を最小化
- [[HAI-LLM]] — DeepSeek LLM の訓練に用いた効率的分散学習フレームワーク
- [[MegaMoE]] — DeepSeek-V4 の MoE 実行カーネル。ウェーブベース融合でエキスパート計算と通信をオーバーラップ
- [[Algirdas Avizienis]] — フォールトトレランス概念の創始者(1967 JPL)・UCLA 教授・[[IFIP WG 10.4]] 創設議長
- [[Jean-Claude Laprie]] — [[LAAS-CNRS]] 研究者・ディペンダビリティ用語の定式化中心人物・[[IFIP WG 10.4]] 議長(1986–1995)
- [[Brian Randell]] — Newcastle 大学教授・リカバリブロック(recovery block)の発明者・ソフトウェア耐障害性先駆者
- [[Carl Landwehr]] — NRL/NSF コンピュータセキュリティ研究者・[[IFIP WG 11.3]] 創設議長
- [[LAAS-CNRS]] — フランス CNRS 傘下のシステム解析・耐障害コンピューティング研究機関(トゥールーズ)
- [[IFIP WG 10.4]] — ディペンダブルコンピューティングとフォールトトレランスに関する IFIP 国際作業グループ
## Transformer / GPT 系
### Person
- [[Ashish Vaswani]] — Transformer 筆頭著者(NeurIPS 2017)。Google Brain → Adept AI → Essential AI
- [[Noam Shazeer]] — Scaled Dot-Product Attention 提案。Google Brain → Character AI 共同創設者
- [[Aidan Gomez]] — Transformer 共著者(NeurIPS 2017)。トロント大学 → Cohere 共同創設者
- [[Illia Polosukhin]] — Transformer 共著者(NeurIPS 2017)。Google Research → NEAR Protocol 共同創設者
- [[Łukasz Kaiser]] — Transformer 共著者・tensor2tensor 実装。Google Brain
- [[Jakob Uszkoreit]] — Transformer 共著者。Google Research → Inceptive Nucleics 共同創設者
- [[Niki Parmar]] — Transformer 共著者。Google Research
- [[Llion Jones]] — Transformer 共著者。Google Research → Sakana AI 共同創設者
- [[Alec Radford]] — GPT-1/GPT-2 筆頭著者、GPT-3 共著者。OpenAI 研究者
- [[Ilya Sutskever]] — OpenAI 共同創設者・元チーフサイエンティスト。GPT-1/2/3 共著者
- [[Karthik Narasimhan]] — GPT-1 共著者。OpenAI → Princeton
- [[Tim Salimans]] — GPT-1 共著者。OpenAI
- [[Dario Amodei]] — GPT-2/3 共著者。OpenAI → Anthropic 共同創設者・CEO
- [[Jeffrey Wu]] — GPT-2 共同第一著者、InstructGPT 共著者。OpenAI
- [[Rewon Child]] — GPT-2/3 共著者、Sparse Transformer。OpenAI
- [[Tom Brown]] — GPT-3 筆頭著者。OpenAI
- [[Jared Kaplan]] — GPT-3 共著者、スケーリング則研究。Johns Hopkins / OpenAI → Anthropic 共同創設者
### Organization
- [[Google Brain]] — Google の AI 研究部門(2023 年に DeepMind と統合し Google DeepMind に)。Transformer の主要開発拠点
- [[OpenAI]] — AI 研究組織(2015 年設立)。GPT シリーズ・DALL·E・ChatGPT の開発元
### Product / System
- [[GPT-2]] — OpenAI の 1.5B パラメータ言語モデル。ゼロショットで 8 データセット中 7 で SOTA
- [[GPT-3]] — OpenAI の 175B パラメータ言語モデル。文脈内学習を大規模に実証(NeurIPS 2020)
### Dataset
- [[WebText]] — OpenAI が構築した 40GB ウェブテキストデータセット(Reddit 3 karma 以上でフィルタ)
## CoNEXT 2026 / ChainScope
## ICML 2026 / SPRINT
### Person
- [[Longlong Xu]] — [[Tsinghua University]] NetManAIOps グループ。SPRINT の第一著者(person)
- [[Zeyan Li]] — [[ByteDance]](Beijing)所属。SPRINT の共著者(person)
### Product
- [[SPRINT]] — 学習不要のプラグアンドプレイ TSFM 推論時拡張フレームワーク。季節パターン複製とトレンド解像度補間で LTSF の精度・効率を同時改善(product)
---
## arXiv 2026 / NexusRCL
### Person
- [[Runzhou Wang]] — [[Nankai University]] の研究者。異種グラフ + 能動学習によるマイクロサービス RCL フレームワーク NexusRCL の筆頭著者(person)
### Product
- [[NexusRCL]] — [[Nankai University]]/[[Tsinghua University]] によるマイクロサービス半教師あり根本原因特定フレームワーク。entity-level 異種グラフ + DBSCAN 能動学習(product)
---
## CoNEXT 2026 / ChainScope
### Person
- [[Ruipeng Hong]] — [[Sun Yat-sen University]] の研究者。eBPF 非侵襲型分散トレーシングシステム ChainScope(CoNEXT 2026)の第一著者(person)
- [[Gabriele Castellano]] — [[Huawei Technologies]](フランス)の研究者。ChainScope の共著者(person)
- [[Massimo Gallo]] — [[Huawei Technologies]](フランス)の研究者。ChainScope の共著者(person)
- [[Zexin Wang]] — CNIC CAS/UCAS の研究者。AgentOps サーベイ論文の第一著者(person)
- [[David Lo]] — Singapore Management University 教授・IEEE Fellow。ソフトウェア工学専門。AgentOps サーベイ共著者(person)
- [[Yintong Huo]] — Singapore Management University の研究者。AgentOps サーベイ共著者(person)
- [[XWind]] — [[Microsoft]] が開発した再生可能エネルギーファーム向け LLM 推論クロスサイトルーター。[[AI Greenferencing]] の制御プレーンを担う(product)
- [[Debopam Bhattacherjee]] — [[Microsoft]] リサーチャー。[[AI Greenferencing]] と [[XWind]] の提唱者(person)
- [[Gaogang Xie]] — CNIC/CAS・Hangzhou Institute for Advanced Study UCAS の研究者。[[UModel]] 共著者(person)
- [[UModel]] — CNIC/CAS・UCAS・[[Alibaba Cloud]] 共同開発のエージェント対応オブザーバビリティデータモデリングフレームワーク(product)
- [[Jason Wei]] — [[Google Brain]] / Google Research 所属の研究者。[[Chain-of-Thought Prompting]] の筆頭著者(person)
- [[Denny Zhou]] — [[Google Brain]] 所属の研究者。[[Chain-of-Thought Prompting]] の共著者(person)
- [[Muhammad Usman]] — [[Karlstad University]] ポスドク研究員。クラウド・エッジコンピューティング・オブザーバビリティ研究者(person)
- [[Simone Ferlin]] — [[Red Hat]] シニア性能エンジニア / 元 [[Karlstad University]] 研究員。MPTCP・輻輳制御の専門家(person)
- [[Anna Brunstrom]] — [[Karlstad University]] 教授・分散システム研究グループ研究マネージャー。IETF RMCAT WG 共同議長(person)
- [[Javid Taheri]] — [[Karlstad University]] 教授。クラウド・SDN 最適化研究者(person)
- [[Karlstad University]] — スウェーデン・カールスタードの大学。分散システム・通信研究グループが 5G・エッジ・オブザーバビリティを研究(organization)
## Ibidunmoye et al. 2015 / CSUR PADBI Survey
### Person
- [[Olumuyiwa Ibidunmoye]] — [[Umeå University]] の研究者。PADBI 分野の体系的サーベイ(ACM CSUR 2015)の第一著者(person)
- [[Francisco Hernández-Rodriguez]] — [[Umeå University]] の研究者。PADBI サーベイの共著者(person)
- [[Erik Elmroth]] — [[Umeå University]] 教授。クラウドコンピューティング・自律システムの専門家。PADBI サーベイのシニア著者(person)
### Organization
- [[Umeå University]] — スウェーデン北部の大学。クラウドコンピューティング・分散システム管理を専門とする研究グループを擁する(organization)
## Zhao et al. 2023 (ISSRE) / Wu et al. 2023 (ICSE-SEIP) — 変更起因インシデント
### Person
- [[Yujin Zhao]] — [[Peking University]]/[[Alibaba Group]] 共同所属。変更起因インシデントライフサイクル論文(ISSRE 2023)の第一著者(person)
- [[Ling Jiang]] — [[Alibaba Group]] エンジニア。変更起因インシデントライフサイクル論文の共著者(person)
- [[Ye Tao]] — [[Peking University]] の研究者。変更起因インシデントライフサイクル論文の共著者(person)
- [[Songlin Zhang]] — [[Alibaba Group]] エンジニア。変更起因インシデントライフサイクル論文の共著者(person)
- [[Changlong Wu]] — [[Alibaba Group]] エンジニア。変更起因インシデントライフサイクル論文の共著者(person)
- [[Yifan Wu]] — [[Peking University]] の研究者。ISSRE 2023 論文の共著者かつ ICSE-SEIP 2023 論文の第一著者(person)
- [[Zhonghai Wu]] — [[Peking University]] 教授。変更起因インシデントライフサイクル論文のシニア著者(person)
- [[Bingxu Chai]] — [[Ant Group]] エンジニア。Ant Group 変更起因インシデント実証研究(ICSE-SEIP 2023)の共著者(person)
- [[Bingchang Liu]] — [[Ant Group]] エンジニア・研究者。ICSE-SEIP 2023 論文の共著者(person)
- [[Jianguo Li]] — [[Ant Group]] シニアエンジニア。ICSE-SEIP 2023 論文の責任著者(person)
- [[Yong Yang]] — [[Peking University]] の研究者。ICSE-SEIP 2023 論文の共著者(person)
- [[Wei Jiang]] — [[Ant Group]] エンジニア。ICSE-SEIP 2023 論文の共著者(person)
- [[Vaibhav Ganatra]] — [[Microsoft]] India 所属。クラウドモニタリングのミス検知タクソノミを分析した [[@2023__ESEC-FSE__Detection Is Better Than Cure - A Cloud Incidents Perspective]](ESEC/FSE 2023)の筆頭著者(person / aiops)
- [[Yu Kang]] — [[Microsoft]] China 所属。[[@2023__ESEC-FSE__Detection Is Better Than Cure - A Cloud Incidents Perspective]](ESEC/FSE 2023)の共著者(person / aiops)
- [[Anjaly Parayil]] — [[Microsoft]] India 所属。[[@2023__ESEC-FSE__Detection Is Better Than Cure - A Cloud Incidents Perspective]](ESEC/FSE 2023)の共著者(person / aiops)
- [[Luan Pham]] — [[RMIT University]] の研究者。マイクロサービス因果推論ベース RCA の包括評価と [[@2024__FSE__BARO - Robust Root Cause Analysis for Microservices via Multivariate Bayesian Online Change Point Detection|BARO]]・[[RCAEval]] の開発者(person / aiops)
- [[Huong Ha]] — [[RMIT University]] の研究者。[[Luan Pham]] の共著者としてマイクロサービス RCA を研究(person)
- [[Hongyu Zhang]] — [[Chongqing University]] の研究者。ソフトウェア工学・AIOps・マイクロサービス RCA を研究(person)
- [[RMIT University]] — オーストラリア・メルボルンの大学。マイクロサービス RCA 研究の拠点(organization)
- [[Chongqing University]] — 中国・重慶の大学。AIOps・ソフトウェア工学研究で知られる(organization)
- [[RCAEval]] — [[Luan Pham]] ら開発のマイクロサービス因果推論ベース RCA 評価フレームワーク(repository / aiops)
## Cohen+ OSDI 2004 — TAN ベース SLO 違反分類
### Person
- [[Ira Cohen]] — [[HP Labs]] 所属。TAN ベース SLO 違反分類の筆頭著者(OSDI 2004)
- [[Moises Goldszmidt]] — [[HP Labs]] 所属。確率的グラフィカルモデルの専門家。TAN ベース SLO 違反分類の共著者(OSDI 2004)
- [[Terence Kelly]] — [[HP Labs]] 所属。TAN ベース SLO 違反分類の共著者(OSDI 2004)
- [[Julie Symons]] — [[HP Labs]] 所属。TAN ベース SLO 違反分類の共著者(OSDI 2004)
- [[Jeffrey S. Chase]] — [[Duke University]] コンピュータサイエンス学科所属。TAN ベース SLO 違反分類の共著者(OSDI 2004)
### Organization
- [[HP Labs]] — Hewlett-Packard の企業研究部門(Palo Alto, CA)。[[Ira Cohen]] ら 4 名が TAN ベース SLO 違反分類を発表(OSDI 2004)
- [[Duke University]] — 米国ノースカロライナ州ダーラムの研究大学。[[Jeffrey S. Chase]] が HP Labs との共同研究に参加(OSDI 2004)
## Chronix FAST '17 — Lautenschlager et al.
### Person
- [[Florian Lautenschlager]] — [[QAware GmbH]] 所属。ドメイン固有 TSDB [[Chronix]] の筆頭開発者(FAST '17)(person)
- [[Michael Philippsen]] — [[Friedrich-Alexander-Universität Erlangen-Nürnberg]] プログラミングシステムグループの教授。[[Chronix]] 共同開発者(FAST '17)(person)
- [[Andreas Kumlehn]] — [[Friedrich-Alexander-Universität Erlangen-Nürnberg]] 所属研究者。[[Chronix]] 共同開発者(FAST '17)(person)
- [[Josef Adersberger]] — [[QAware GmbH]] 共同創業者・CTO。[[Chronix]] 共同開発者(FAST '17)(person)
### Organization
- [[QAware GmbH]] — ドイツ・ミュンヘン拠点のソフトウェアエンジニアリング会社。[[Chronix]] の産業側開発拠点(organization)
- [[Friedrich-Alexander-Universität Erlangen-Nürnberg]] — ドイツ・バイエルン州の研究大学(FAU)。[[Michael Philippsen]] プログラミングシステムグループが [[QAware GmbH]] と共同で [[Chronix]] を開発(organization)
### Product
- [[Chronix]] — 運用データ異常検知に特化したオープンソースのドメイン固有時系列データベース。DDC・汎用データモデル・ビルトイン解析・コミッショニング方法論を特徴とする(product / time-series / database)
- [[Lindorm TSDB]] — [[Alibaba Group]] が開発するクラウドネイティブ分散 TSDB。共有なし + 共有ストレージのハイブリッドアーキテクチャ・TSM・Seriescache・前処理ダウンサンプリング・Lindorm ML を特徴とする(product / time-series / database)
## Lindorm TSDB PVLDB 2023 — Shen et al.
### Person
- [[Feifei Li]] — [[Alibaba Group]] データベース部門。大規模分散データベース・時系列データ管理専門。[[Lindorm TSDB]] 主要設計者(PVLDB 2023)(person)
### Organization
- [[Zhejiang University]] — 中国・杭州市の重点大学(985 工程)。[[Lindorm TSDB]] 論文に Chunhui Shen が参加(organization)
### AIM Survey (JNCA 2024) — added 2026-06-14
- [[Qingyang Yu]] — [[Tsinghua University]] 計算機科学技術学科の研究者。[[Dan Pei]] グループ。AIM サーベイ(JNCA 2024)筆頭著者(person / aiops / survey)
### Distributed Message Broker Survey (arXiv 2017) — added 2026-06-14
- [[Vineet John]] — [[University of Waterloo]] 研究者(2017 年時点)、Flotilla ベンチ作者(person / messaging)
- [[Xia Liu]] — [[University of Waterloo]] 研究者(2017 年時点)、Kafka vs AMQP 比較共著(person / messaging)
- [[University of Waterloo]] — カナダ・オンタリオ州の公立研究大学(organization / university)
- [[Apache Kafka]] — [[LinkedIn]] 発の分散ストリーミングプラットフォーム。ログ処理用途、高スループット・低レイテンシ志向(product / messaging)
- [[AMQP]] — Advanced Message Queuing Protocol。金融取引用途由来、信頼性・スケーラビリティ・管理性高保証(product / messaging / protocol)
- [[RabbitMQ]] — VMware による Erlang 製 [[AMQP]] 人気実装(product / messaging)
### TTMPred / ISSRE 2021 — added 2026-06-14
- [[Weijing Wang]] — [[Tianjin University]] の研究者。インシデント TTM 予測の最初の深層学習アプローチ TTMPred を主導(person / aiops / incident-management)
- [[Tianjin University]] — 中国・天津市の国立研究大学(天津大学)。College of Intelligence and Computing が AI・ソフトウェア工学研究を担う(organization / university)
### Li+ ISSRE 2022 — added 2026-06-14
- [[Xiaoyun Li]] — [[Sun Yat-sen University]] 所属(2022 年時点)。[[@2022__ISSRE__Going through the Life Cycle of Faults in Clouds - Guidelines on Fault Handling]] の共同第一著者([[Guangba Yu]] と同等貢献)。Email:
[email protected](person / cloud-reliability / empirical-study)
- [[Hongyang Chen]] — [[Sun Yat-sen University]] 所属(2022 年時点)。Li+ ISSRE 2022 の共著者。Email:
[email protected](person / cloud-reliability)
- [[Zhekang Chen]] — [[Bizseer]] 所属。Li+ ISSRE 2022 の産業界共著者。Email:
[email protected](person / cloud-reliability / industry)
### DeepIP / ASE 2020 — added 2026-06-15
- [[Junjie Chen]] — [[Tianjin University]] College of Intelligence and Computing 助教。[[DeepIP]] 論文筆頭著者(person / aiops / incident-management)
- [[Shu Zhang]] — [[Microsoft]] Research Beijing 所属(2020)。DeepIP 共著者(person / aiops)
- [[Xiaoting He]] — [[Microsoft]] Research Beijing 所属(2020)。DeepIP 共著者(person / aiops)
- [[Dan Hao]] — [[Peking University]] 所属。ソフトウェア工学・AIOps 研究者。DeepIP 共著者(person / software-engineering / aiops)
- [[Feng Gao]] — [[Microsoft Azure]] Redmond 所属。DeepIP 共著者(person / azure)
- [[Zhangwei Xu]] — [[Microsoft Azure]] Redmond 所属。DeepIP 共著者(person / azure)
- [[Yingnong Dang]] — [[Microsoft Azure]] Redmond 所属。DeepIP 共著者(person / azure)
- [[University of Newcastle]] — オーストラリア Callaghan 所在の公立研究大学。[[Hongyu Zhang]] が 2020〜2021 年時点で所属(organization / university)
- [[DeepIP]] — Microsoft の attention 付き CNN + 関連 incident 取り込みベースのインシデント優先順位付け手法。AUC 平均 0.808 で bug severity prediction 流用を 18 全システムで上回る(product / aiops / incident-management)
- [[Qinwei Ma]] — [[Tsinghua University]] の研究者。時系列予測の汎ドメインアーキテクチャ限界のポジションペーパー筆頭著者(person / time-series)
- [[Jingzhe Shi]] — [[Tsinghua University]] の研究者。同ポジションペーパー共著者(person / time-series)
- [[Jiahao Qiu]] — [[Princeton University]] の研究者。TFB ベンチマーク共著者(person / time-series)
- [[Zaiwen Yang]] — [[Tsinghua University]] の研究者。同ポジションペーパー共著者(person / time-series)
### Time Series Reasoning + Video Grounding batch — added 2026-06-15
- [[Jiahao Wang]] — [[University of Science and Technology of China]] の研究者。TimeReasoner(WSDM 2026)の第 2 著者(person / time-series / llm)
- [[Daoyu Wang]] — [[University of Science and Technology of China]] の研究者。TimeReasoner(WSDM 2026)の第 3 著者(person / time-series / llm)
- [[Tong Guan]] — [[Griffith University]] / [[Zhejiang University]] の研究者。[[@2026__ICLR2026__TimeOmni-1 - Incentivizing Complex Reasoning with Time Series in Large Language Models|TimeOmni-1]] と TSR-Suite の筆頭著者(person / time-series / llm / reasoning)
- [[Qin Jin]] — [[Renmin University of China]] AIM3 Lab の教授。Time-R1(NeurIPS 2025)の対応著者(person / video-understanding / vlm)
- [[MiLM Plus]] — [[Xiaomi]] Inc. の AI 研究組織。LLM・VLM の研究開発(organization / industry / llm)
- [[Xiaomi]] — 中国の総合家電・スマートデバイス企業。MiLM Plus を擁する(organization / industry)
- [[Xiaohan Zhang]] — [[University of Science and Technology of China]] の研究者。AlphaCast の筆頭著者(person / time-series / llm)
- [[Tian Gao]] — [[University of Science and Technology of China]] の研究者。AlphaCast の第 2 著者(person / time-series / llm)
- [[Winnie Chow]] — [[Stanford University]] の研究者(Apple インターン)。Towards Time-Series Reasoning with LLMs(NeurIPS 2024 Workshop)の筆頭著者(person / time-series / llm)
- [[Lauren Gardiner]] — [[Apple]] の研究者。同論文共著者(person / time-series / llm)
- [[Haraldur T. Hallgrimsson]] — [[Apple]] の研究者。同論文共著者(person / time-series / llm)
- [[Maxwell A. Xu]] — [[University of Illinois Urbana-Champaign]] の研究者(Apple インターン)。同論文共著者(person / time-series / llm)
- [[Shirley You Ren]] — [[Apple]] の研究者。同論文共著者(person / time-series / llm)
- [[Apple]] — アメリカのテクノロジー企業。機械学習・時系列解析の研究グループを擁する(organization / industry)
### NSDI'24 Acme paper batch — added 2026-06-15
- [[Dahua Lin]] — [[Shanghai AI Laboratory]] / [[The Chinese University of Hong Kong]] の研究者。[[InternLM]] 系 LLM 開発を主導し [[@2024__NSDI__Characterization of Large Language Model Development in the Datacenter]] 共著(person / llm-systems)
- [[Yonggang Wen]] — [[Nanyang Technological University]] 教授。AI システム・データセンターワークロード分析の共著者(person / ai-systems)
- [[Nanyang Technological University]] — シンガポールの研究大学。S-Lab を含む AI システム研究拠点で Acme LLM クラスタ特性化研究の共同所属組織(organization / academic)
- [[SenseTime Research]] — 中国 [[SenseTime Research|SenseTime]] の研究組織。[[Helios]] GPU データセンタートレース(2020)と [[Acme]] 特性化研究に参加(organization / industry / ai)
- [[InternLM]] — [[Shanghai AI Laboratory]] が [[Acme]] で開発する LLM シリーズ(7B〜123B、transformer decoder-only)(product / llm)
- [[AcmeTrace]] — [[Acme]] の Seren/Kalos 2 クラスタ・2023-03〜2023-08 のスケジューラログ・ハードウェア監視・ランタイムログ・プロファイリング 4 系統トレースの公開版(dataset / gpu-cluster)
- [[Yuting Jiang]] — [[Microsoft Research]] の研究者。SuperBench の equal contribution 著者(person / ai-systems)
- [[Ziyue Yang]] — [[Microsoft Research]] の研究者。SuperBench の equal contribution 著者(person / ai-systems)
- [[Lei Qu]] — [[Microsoft Research]] の研究者。SuperBench 共著(person / ai-systems)
- [[Yongqiang Xiong]] — [[Microsoft Research]] Asia の研究者。Systems and Networking Research Group。SuperBench シニア共著(person / ai-systems)
- [[Lidong Zhou]] — [[Microsoft Research]] の研究者。グレイ障害概念(Huang+ 2017)提唱者の一人で SuperBench シニア共著(person / distributed-systems / reliability)
- [[SuperBench]] — [[Microsoft]] が Azure に 2 年以上デプロイしている GPU クラスタ向けプロアクティブ検証システム。Validator(CDF 類似度クラスタリング)+ Selector(Cox-Time + 貪欲法)+ ベンチマーク群。OSS は microsoft/superbenchmark(product / aiops / gpu)
- [[CIRCA]] — [[Mingjie Li]]・[[Dan Pei]] ほか([[Tsinghua University]] / BizSeer)が KDD 2022 で提案した因果推論ベース RCA。Pearl の Causal Hierarchy で介入認識(IR)タスクとして RCA を定式化し、構造グラフ + 回帰仮説検定(RHT) + 子孫調整で Oracle DB 99 件で AC@1=0.404(ベースライン +25%)。コード: NetManAIOps/CIRCA(product / aiops / rca / causal-inference)
- [[RCD]] — [[Azam Ikram]] ほか([[Purdue University]] / [[Adobe Research]])が NeurIPS 2022 で提案した因果推論ベース RCA。障害を soft intervention とモデル化し、F-NODE 近傍局所学習 + 階層分割統治 Ψ-PC で 500 ノード 22 秒のスケーラビリティ。コールグラフ・パラメトリック仮定不要。コード: azamikram/rcd(product / aiops / rca / causal-inference / microservices)
- [[Kanglin Yin]] — [[Tsinghua University]] の研究者。CIRCA(KDD 2022)共著(person / aiops / rca)
- [[Xiaohui Nie]] — BizSeer の研究者。CIRCA(KDD 2022)共著(person / aiops / rca)
- [[Wenchi Zhang]] — BizSeer の研究者。CIRCA(KDD 2022)共著(person / aiops / rca)
- [[Kaixin Sui]] — BizSeer の研究者。CIRCA(KDD 2022)共著(person / aiops / rca)
- [[Azam Ikram]] — [[Purdue University]] の研究者。RCD(NeurIPS 2022)第一著者。マイクロサービス因果推論ベース RCA を扱う(person / aiops / rca / causal-inference)
- [[Saurabh Bagchi]] — [[Purdue University]] の教授。分散システム・信頼性研究。RCD(NeurIPS 2022)共著(person / distributed-systems / reliability)
- [[Murat Kocaoglu]] — [[Purdue University]] の研究者。因果推論・因果発見が専門。RCD(NeurIPS 2022)共著(person / causal-inference / machine-learning)
- [[Sarthak Chakraborty]] — [[Adobe Research]] の研究者。RCD(NeurIPS 2022)共著、マイクロサービス因果 RCA を扱う(person / aiops / rca)
- [[Purdue University]] — 米国インディアナ州の研究大学。Saurabh Bagchi・Murat Kocaoglu・Azam Ikram が RCD(NeurIPS 2022)を発表(organization / academia / us)
- [[Adobe Research]] — Adobe の研究部門。Sarthak Chakraborty ほかが RCD(NeurIPS 2022)を共同発表(organization / industry-research)
- [[Xiaoming Shi]] — [[Xiaohongshu Inc]] 所属。Time-MoE(arXiv:2409.16040、ICLR 2025)の均等貢献第一著者・プロジェクトリード(♠)(person / time-series / scaling)
- [[Shiyu Wang]] — [[Xiaohongshu Inc]] 所属。Time-MoE の均等貢献第一著者・プロジェクトリード(♠)(person / time-series / scaling)
- [[Yuqi Nie]] — [[Princeton University]] 所属。Time-MoE の均等貢献第一著者。PatchTST 共著者でもある(person / time-series / transformer)
- [[Xiaohongshu Inc]] — 中国のソーシャルコマース・コンテンツプラットフォーム企業(小红書)。Time-MoE(ICLR 2025)の主要所属機関(organization / industry / time-series)
- [[Griffith University]] — オーストラリア・クイーンズランド州の大学。[[Ming Jin]]・[[Shirui Pan]] の所属機関。Time-MoE・TimeOmni-1 の共著機関(organization / university / time-series)
- [[Time-300B]] — Time-MoE の事前学習用大規模公開時系列データコレクション。9 ドメイン・48M 系列・309B 観測点(dataset / time-series)
- [[Hao Zhang]] — [[National University of Singapore]] School of Computing 所属。TKDE 2015 In-Memory Big Data Survey 第一著者。MemepiC・UVMM・Anti-Caching 系列の共著者(person / database)
- [[Gang Chen]] — [[Zhejiang University]] 計算機科学学院教授、IEEE Member。NUS DB Group と長年共同(person / database)
- [[Beng Chin Ooi]] — [[National University of Singapore]] 教授、IEEE Fellow。NUS Database Group の中核(person / database)
- [[Kian-Lee Tan]] — [[National University of Singapore]] 教授、IEEE Member。クエリ処理・索引・オンライン集約(person / database)
- [[Meihui Zhang]] — [[Singapore University of Technology and Design]] ISTD Pillar(後に BIT 異動)、IEEE Member(person / database)
- [[Singapore University of Technology and Design]] — シンガポールの工科系研究大学(2009 設立、MIT 共同設立支援)。Information Systems Technology and Design Pillar(organization / academia)
- [[Cristian Challu]] — [[Nixtla]] の研究者。TimeGPT-1(arXiv:2310.03589)共著者。N-HiTS(AAAI 2023)筆頭著者(person / time-series / foundation-model)
- [[Max Mergenthaler-Canseco]] — [[Nixtla]] の研究者。TimeGPT-1(arXiv:2310.03589)共著者(person / time-series / foundation-model)
- [[Nixtla]] — サンフランシスコ拠点の時系列予測スタートアップ。TimeGPT-1 を開発し REST API と Python SDK で提供。StatsForecast・NeuralForecast のオープンソースも提供(organization / time-series)
- [[Nicole Forsgren]] — [[DORA]] 共同創設者・研究者。"Accelerate"(2018)・[[SPACE]] フレームワーク(2021)・"Frictionless"(2026、[[Abi Noda]] と共著)の著者。SREcon26 で DX を信頼性システム特性として論じた(person / sre / developer-experience)
- [[Abi Noda]] — 開発者体験の専門家。[[Nicole Forsgren]] と "Frictionless"(2026)共著(person / developer-experience)
- [[Jamie Wilkinson]] — [[Google]] SRE。SREcon18 Asia でシンプトムベースドアラーティングと SLO バーンレートアラートを体系化(person / sre / slo)
- [[Annie Zhou]] — [[Databricks]] ストレージプラットフォームチームのエンジニア。[[Sophie Zhang (Databricks)]] と SREcon26 Americas で AI支援 DB デバッグシステムを発表(person / sre / databricks)
- [[Sophie Zhang (Databricks)]] — [[Databricks]] ストレージプラットフォームチームのエンジニア。[[Annie Zhou]] と SREcon26 Americas で [[Storax]] の安全基盤設計を発表(person / sre / databricks)
- [[Databricks]] — データ・AI プラットフォーム企業。Apache Spark・Delta Lake・MLflow・Unity Catalog の開発元。[[Storax]] で AI支援 DB O&M を本番化(organization / industry / data-platform)
- [[Storax]] — [[Databricks]] 内部の AI デバッグツールを支えるストレージプラットフォームサービス。セントラルファースト・シャーデッドアーキテクチャ、細粒度 AC、統一ツールインターフェース、Temporal 承認ゲートを持つ(product / sre / aiops / database)
- [[Tianyi Yang]] — [[The Chinese University of Hong Kong]] Lyu 研究室。Yang+ DSN2022 アラートアンチパターン筆頭(person / aiops)
- [[Jiacheng Shen]] — CUHK Lyu 研究室。Yang+ DSN2022 第 2 著者(person / aiops)
- [[Yuxin Su]] — [[Sun Yat-sen University]]。Yang+ DSN2022 責任著者(person / aiops)
- [[Xiaoxue Ren]] — CUHK。Yang+ DSN2022 共著(person / aiops)
- [[Jinxi Kuang]] — CUHK Lyu 研究室。Kuang+ ICSE-SEIP2024 [[COLA]] 筆頭(person / aiops)
- [[Jinyang Liu]] — CUHK Lyu 研究室。Kuang+ ICSE-SEIP2024 第 2 著者、Logzip/iPACK 共著(person / aiops)
- [[Jiazhen Gu]] — CUHK Lyu 研究室。Kuang+ ICSE-SEIP2024 責任著者(person / aiops)
- [[Lan Yu]] — [[Huawei Cloud]] Computing and Networking Innovation Lab。Kuang+ ICSE-SEIP2024 共著(person / cloud-reliability)
- [[Rui Tan]] — [[Huawei Cloud]] Computing and Networking Innovation Lab。Kuang+ ICSE-SEIP2024 共著(person / cloud-reliability)
- [[Akanksha Singal]] — [[IBM Research]] India + [[IIIT Delhi]]。Singal+ arXiv2025 [[KIMetrix]] 筆頭(person / aiops / microservices)
- [[Kaustabha Ray]] — [[IBM Research]] India。Singal+ arXiv2025 共著(person / aiops)
- [[Divya Pathak]] — [[IBM Research]] India。Singal+ arXiv2025 共著(person / aiops)
- [[Felix George]] — [[IBM Research]] India。Singal+ arXiv2025 共著(person / aiops)
- [[Mudit Verma]] — [[IBM Research]] India。Singal+ arXiv2025 共著(person / aiops)
- [[Pratibha Moogi]] — [[IBM Research]] India。Singal+ arXiv2025 シニア共著(person / aiops)
- [[IIIT Delhi]] — Indraprastha Institute of Information Technology Delhi。[[Akanksha Singal]] 所属(organization / university / india)
- [[Derek Lin]] — [[Pivotal Software]] Palo Alto。Lin+ KDD2014 アラート/インシデントクラスタリング筆頭(person / aiops / clustering)
- [[Rashmi Raghu]] — [[Pivotal Software]] Palo Alto。Lin+ KDD2014 共著(person / aiops)
- [[Vivek Ramamurthy]] — [[Pivotal Software]] Palo Alto。Lin+ KDD2014 共著(person / aiops)
- [[Jin Yu]] — [[Pivotal Software]] Melbourne。Lin+ KDD2014 共著、NMF + KD-tree 部担当(person / aiops)
- [[Regunathan Radhakrishnan]] — [[Pivotal Software]] Palo Alto。Lin+ KDD2014 共著(person / aiops)
- [[Joseph Fernandez]] — [[Visa Inc]] Foster City。Lin+ KDD2014 共著、Pivotal とのデータ提供連携(person / aiops / industry)
- [[Jim Thomson]] — [[Pivotal Software]] Cloud R&D Product Lead。脆弱性バジェット・フィーチャーフレッシュネスを David Laing と共に提唱(SREcon19 Americas)(person / sre / error-budget)
- [[David Laing]] — [[Pivotal Software]] Cloud R&D Engineering Lead、CloudOpsEU 元責任者。脆弱性バジェットの実測グラフを公開(SREcon19 Americas)(person / sre / error-budget)
- [[Pivotal Software]] — Palo Alto + Melbourne。Greenplum DB + MADlib の MPP データ分析企業。Cloud R&D チームが SRE 実践(脆弱性バジェット)を推進(organization / data-analytics / sre)
- [[Visa Inc]] — Foster City, CA。決済ネットワーク事業者。Lin+ KDD2014 のデータ提供側(organization / finance)
- [[Yujun Chen]] — [[Beihang University]] Beijing + [[Microsoft Research]] インターン。Chen+ WWW2019 [[AirAlert]] 筆頭(person / aiops / outage-prediction)
- [[Hang Dong]] — [[Microsoft Research]] Beijing。Chen+ WWW2019 共著(person / aiops)
- [[Fotios Voutsas]] — [[Netdata]] Inc. エンジニア。Voutsas+ JCC2023 アラートフィルタリング二値分類の筆頭著者(person / aiops / alert-management)
- [[John Violos]] — [[École de Technologie Supérieure]](ÉTS, モントリオール)ソフトウェア・IT エンジニアリング学科。Voutsas+ JCC2023 共著(person / aiops / alert-management)
- [[Aris Leivadeas]] — [[École de Technologie Supérieure]](ÉTS, モントリオール)ソフトウェア・IT エンジニアリング学科。NSERC 助成 RGPIN-2019-05250 主要研究者。Voutsas+ JCC2023 共著(person / aiops / alert-management)
- [[Netdata]] — リアルタイムモニタリング・可視化ツールの開発企業(Netdata Inc.)。JCC2023 アラートフィルタリング研究でデータ提供・共同研究パートナー(organization / cloud-monitoring / aiops)
- [[École de Technologie Supérieure]] — カナダ・モントリオールの工科大学(ÉTS)。Violos・Leivadeas の所属。クラウドモニタリング研究を実施(organization / university / canada)
- [[Yiru Chen]] — [[Fudan University]] 大学院生。Chen+ ASE2023 (DyAlert) 第一著者。アラートリンク予測の動的グラフ手法(person / aiops)
- [[Yuqun Zhang]] — [[Southern University of Science and Technology]](SUSTech)助教。Zeng+ ICSE-SEIP2023 (TraceArk) 共同筆頭著者(person / software-engineering)
- [[Yuan Yuan]] — [[National University of Defense Technology]] 大学院生。Yuan+ ISSRE2024 (SuperAgg) 第一著者(person / aiops / hpc)
- [[Tongqing Zhou]] — NUDT 准教授。Yuan+ ISSRE2024 (SuperAgg) 共著(person / aiops / hpc)
- [[Karan Bhukar]] — [[IBM Research]] India。Bhukar+ ICSE-SEIP2024 (Dynamic-X-Y) 第一著者(person / aiops)
- [[Pooja Aggarwal]] — [[IBM Research]] India。Bhukar+ ICSE-SEIP2024 共著。アラート抑制・AIOps 研究(person / aiops)
- [[Yongqian Sun]] — [[Nankai University]] 准教授。Yuan+ ISSRE2024 (SuperAgg) 共著、アラート管理研究(person / aiops)
- [[Zhen Dong]] — [[Fudan University]]。Chen+ ASE2023 (DyAlert) 共著(person / software-engineering)
- [[Xin Peng]] — [[Fudan University]] 教授。Chen+ ASE2023 (DyAlert) 共著(person / software-engineering)
- [[Shubham Agarwal]] — [[Adobe Research]] India。Chakraborty+ ESRO 共著(person / aiops)
- [[Shaddy Garg]] — [[Adobe]] India。Chakraborty+ ESRO 共著(person / aiops)
- [[Shiv Saini]] — [[Adobe Research]] India。Chakraborty+ ESRO 対応著者(person / aiops)
- [[ESRO]] — Chakraborty+ arXiv2023 提案の経験ベース診断サービス。CK グラフでアラート+障害レポートを統合、Random Forest クラスタ予測で top-1 62%(product / aiops)
- [[AlertRank]] — Zhao+ ISSRE2020 提案のアラート重要度ランキング手法。XGBoost + 40 次元特徴量で F1=0.89、Resolution Record の TF-IDF 自動ラベル(product / aiops)
- [[AlertRCA]] — Yu+ CCGRID2024 提案のアラートのみ入力 RCA。Alert2Vec + CPGAT + DAGNN で top-1 83.9%、Groot を上回る(product / aiops)
- [[TraceArk]] — Zeng+ ICSE-SEIP2023 提案のアクショナブルアラート手法。ExL + パス粒度集約 + XGBoost フィードバック、Exchange 本番 4 ヶ月で適合率 0.9068(product / aiops)
- [[SuperAgg]] — Yuan+ ISSRE2024 提案の HPC 向けアラート集約。センサ層 4 パターン + Apriori 主従関係の 2 段階階層構造、集約率 99.04%(product / aiops / hpc)
- [[Fudan University]] — 中国・上海の総合大学。Chen+ ASE2023 (DyAlert) の所属(organization / university / china)
- [[IIT Kanpur]] — Indian Institute of Technology Kanpur。ESRO 共著者の所属(organization / university / india)
- [[Stevens Institute of Technology]] — 米国ニュージャージーの工科大学。AlertRank 共著者 Rong Liu の所属(organization / university / usa)
- [[Jiayu Hu]] — [[Tencent]] 所属エンジニア。Harp 第一著者、NSDI 2026(person / networking)
- [[Feng Jin]] — [[Tencent]] 所属エンジニア。Harp 共著者(person / networking)
- [[Kai Zhang]] — [[Fudan University]] 所属。Harp の corresponding author、NSDI 2026(person / networking)
- [[LLMAD]] — Liu+ KDD2025 提案の LLM 直接判定 TSAD。GPT-4 + FastDTW ICL + AnoCoT で KPI/WSD/Yahoo 平均 Best F1=0.759、年間 $65.70(product / aiops / anomaly-detection / llm)
- [[ChatTS]] — Xie+ VLDB2025 提案の初の TS-MLLM。Qwen2.5-14B を完全合成データで SFT、GPT-4o vision を alignment +46% / reasoning +26% で凌駕(product / aiops / llm / time-series / multimodal)
- [[AnoCoT]] — LLMAD で提案された TSAD 専用 CoT。判定ルール + 異常タイプ定義 + アラームレベル定義 + 大域 → 局所 → 再評価の 3 ステップ(product / llm / anomaly-detection)
- [[TSEvol]] — ChatTS の合成 Q&A 生成器。Evol-Instruct を時系列属性プールに拡張、Breadth/Depth/Reasoning/Situation の 4 種進化(product / llm / synthetic-data)
- [[Anomaly Transformer]] — Xu+ 2021 の Transformer 教師なし TSAD。association discrepancy + anomaly-attention 機構(product / anomaly-detection / time-series)
- [[Qwen2.5-14B-Instruct]] — Alibaba 公開の 14B Instruct LLM。ChatTS のベース LLM(product / llm)
- [[Jun Liu (UCAS)]] — [[University of Chinese Academy of Sciences]] の研究者。LLMAD 第一著者(Microsoft インターン)(person)
- [[Jiaxu Qian]] — [[Zhejiang University of Technology]] の研究者。LLMAD 共第一著者(Microsoft インターン)(person)
- [[Xiao He]] — [[ByteDance]] の研究者。ChatTS 共著者、OneShotSTL(PVLDB 2023)主著者(person)
- [[Jianjun Chen]] — [[ByteDance]] の研究者。ChatTS 共著者(person)
- [[Rui Shi]] — [[ByteDance]] の研究者。ChatTS 共著者(person)
- [[University of Chinese Academy of Sciences]] — 中国・北京の研究大学。LLMAD 第一著者 Jun Liu の所属(organization / university / china)
- [[Zhejiang University of Technology]] — 中国・杭州の工科大学。LLMAD 共第一著者 Jiaxu Qian の所属(organization / university / china)
### 2026-06-17 分散深層学習の訓練系基盤論文 14 本一括
#### People
- [[Philipp Moritz]] — [[University of California, Berkeley]]、[[Ray]] 第一著者(person / distributed-systems)
- [[Ion Stoica]] — [[University of California, Berkeley]]、Ray / Spark 共同創設者(person / distributed-systems)(更新)
- [[Mohammad Shoeybi]] — [[NVIDIA]]、[[Megatron-LM]] 第一著者(person / nlp)
- [[Bryan Catanzaro]] — [[NVIDIA]] VP of Applied Deep Learning Research(person / deep-learning)
- [[Deepak Narayanan]] — [[Stanford University]] / [[Microsoft Research]]、PipeDream・PTD-P 第一著者(person / distributed-systems)
- [[Matei Zaharia]] — [[Stanford University]] / [[Databricks]]、Spark / Ray / PipeDream 共著者(person / distributed-systems)
- [[Yanping Huang]] — [[Google Brain]]、GPipe 第一著者(person / deep-learning)
- [[Quoc V. Le]] — [[Google Brain]]、GPipe・AutoML 研究(person / deep-learning)
- [[Jeff Rasley]] — [[Microsoft]]、DeepSpeed 創設開発者(person / distributed-systems)
- [[Samyam Rajbhandari]] — [[Microsoft]]、ZeRO / DeepSpeed コア技術開発者(person / hpc)
- [[Yuxiong He]] — [[Microsoft]] Research Manager、DeepSpeed プロジェクトリーダー(person / deep-learning)
- [[Hanyu Zhao]] — [[Peking University]]、HiveD 第一著者(person / gpu-scheduling)
- [[Vijay Korthikanti]] — [[NVIDIA]]、選択的活性化再計算・シーケンス並列化 第一著者(person / deep-learning)
- [[Houwen Peng]] — [[Microsoft Research]]、FP8-LM 第一著者(person / deep-learning)
- [[Han Hu]] — [[Microsoft Research Asia]]、FP8-LM 共著者(person / computer-vision)
- [[Yanli Zhao]] — [[Meta]]、PyTorch FSDP 第一著者(person / distributed-systems)
- [[Kai Chen (HKUST)]] — [[Hong Kong University of Science and Technology|HKUST]] [[iSING Lab]] 主宰(person / networking)
- [[Sudarsanan Rajasekaran]] — [[MIT]]、Cassini 第一著者(person / networking)
- [[Manya Ghobadi]] — [[MIT]] CSAIL、Cassini 共著者(person / networking)
- [[Aditya Akella]] — [[UT Austin]]、Cassini 共著者(person / networking)
- [[Bohan Zhao]] — [[Tsinghua University]]、FFTrainer 第一著者(person / distributed-systems)
- [[Wei Xu]] — [[Tsinghua University]]、FFTrainer 共著者(person / distributed-systems)
- [[Peng Cheng]] — [[Microsoft Research Asia]]、通信特性論文共著者(person / networking)(更新)
#### Products / Systems
- [[Ray]] — UC Berkeley 発の分散 AI フレームワーク、タスク並列 + アクターモデル統合(product / distributed-framework)(更新)
- [[GPipe]] — Google Brain のマイクロバッチパイプライン並列化ライブラリ(product / pipeline-parallelism)
- [[PipeDream]] — Microsoft Research / Stanford の 1F1B パイプライン並列訓練システム(product / pipeline-parallelism)
- [[DeepSpeed]] — Microsoft の大規模分散訓練ライブラリ(product / distributed-training)(更新)
- [[Megatron-LM]] — NVIDIA の大規模言語モデル訓練フレームワーク(product / distributed-training)(更新)
- [[HiveD]] — Microsoft / Peking University のマルチテナント GPU クラスタスケジューラ(product / gpu-scheduling)
- [[OpenPAI]] — Microsoft のオープンソース AI プラットフォーム、HiveD を組み込み(product / platform)
- [[Cassini]] — MIT のネットワーク対応 ML ジョブスケジューラ(product / gpu-scheduling)
- [[FFTrainer]] — Tsinghua University のゼロオーバーヘッド障害復旧フレームワーク(product / fault-tolerance)
#### Organizations
- [[iSING Lab]] — HKUST Kai Chen 研究室、ネットワーク + AI インフラ研究(organization / research-lab)
- [[University of California, Berkeley]] — Ray・Spark 発祥の地(organization / university)(更新)
#### JANOG56 (2026-06-19)
- [[サイバーエージェント]] — 日本のインターネット企業。Cycloud プライベートクラウドで AI/ML 基盤を運営し、TOP500 国内 15 位達成(organization)
- [[CIU]] — サイバーエージェントグループの IT 推進本部・インフラ部門。AS24284 運用・Cycloud 運営(organization)
- [[小障子 尚太朗]] — サイバーエージェント CIU ネットワークエンジニア。GPU 基盤ネットワーク設計・JANOG52/56 登壇(person / networking)
- [[疋田 紅樹]] — サイバーエージェント CIU ネットワークエンジニア。パッチパネル高密度化・負荷分散チューニング・JANOG56 登壇(person / networking)
- [[Juniper QFX5240]] — 2U 64 ポート 800G スイッチ(Broadcom ASIC / JunOS)。Cycloud GPU インターコネクト Leaf として採用(product / networking)
#### クラスタリング基礎論文 (2026-06-20)
- [[Martin Ester]] — Simon Fraser University 教授。DBSCAN 共同提案者(person / data-mining)
- [[Hans-Peter Kriegel]] — ミュンヘン大学(LMU)教授。DBSCAN・R*-tree 共著者(person / data-mining)
- [[Jörg Sander]] — アルバータ大学教授。DBSCAN・OPTICS・HDBSCAN に貢献(person / data-mining)
- [[Ricardo J.G.B. Campello]] — サンパウロ大学(USP)教授。HDBSCAN 提案者(person / data-mining)
- [[Davoud Moulavi]] — アルバータ大学研究者。HDBSCAN 共著者(person / data-mining)
- [[John Paparrizos]] — Ohio State University 助教(元 Columbia University)。k-Shape 提案者(person / data-mining / time-series)
- [[Luis Gravano]] — Columbia University 教授。k-Shape 共著者(person / databases)
#### マイクロサービス・DB RCA 基礎論文 (2026-06-20)
- [[Myunghwan Kim]] — Stanford University 博士課程(2013 年当時)。MonitorRank 提案者(person / data-mining)
- [[LinkedIn]] — ビジネス SNS。MonitorRank 論文時に 400 超のサービスを運用(organization)
- [[Stanford University]] — スタンフォード大学(organization / university)(更新)
- [[Chunqiu Zeng]] — Florida International University。時間遅れマイニング研究(person / data-mining)
- [[Tao Li]] — Florida International University 教授。テキスト・時系列マイニング(person / data-mining)
- [[Florida International University]] — マイアミの州立大学(organization / university)
- [[Larisa Shwartz]] — IBM T.J. Watson Research Center。運用自動化研究(person / systems)
- [[Genady Grabarnik]] — St. John's University 数学・CS 学部教授(person / mathematics)
- [[Ping Wang]] — Peking University。CloudRanger 筆頭著者(person / systems)
- [[Peking University]] — 北京大学(organization / university)(更新)
- [[Ping Liu]] — Tsinghua University / BNRist。FluxRank・FluxInfer 筆頭著者(person / aiops)
- [[Shenglin Zhang]] — Nankai University。FluxRank・FluxInfer 共著者(person / aiops)(更新)
- [[Jing Zhu (Tsinghua)]] — Tsinghua University。FluxRank 共著者(person / aiops)
- [[Yu Chen (Baidu)]] — Baidu。FluxRank 共著者(person / industry)
- [[Huasong Shan]] — JD.com。ε-Diagnosis 筆頭著者(person / aiops)
- [[JD.com]] — 中国大手 EC プラットフォーム(organization)
- [[Minghua Ma]] — Tsinghua University / Microsoft Research。AutoMAP 共著者(person / aiops)(更新)
- [[Li Wu]] — TU Berlin 博士課程。MicroDiag 筆頭著者(person / aiops)
- [[Johan Tordsson]] — Umeå University 准教授。クラウドリソース管理(person / distributed-systems)
- [[Odej Kao]] — TU Berlin 教授。分散システム・機械学習(person / distributed-systems)
- [[TU Berlin]] — ベルリン工科大学(organization / university)
- [[Umeå University]] — ウメオ大学(organization / university)
- [[Elastisys AB]] — スウェーデンのクラウド自動化企業(organization)
- [[Jasmin Bogatinovski]] — TU Berlin 研究者。AIOps(person / aiops)
- [[Erik Elmroth]] — Umeå University 教授。分散システム(person / distributed-systems)
- [[Guangba Yu]] — Sun Yat-sen University。TS-InvarNet 共著者(person / aiops)(更新)
- [[Sun Yat-sen University]] — 中山大学(organization / university)(更新)
- [[Chenghao Liu]] — Salesforce AI。PyRCA 筆頭著者(person / machine-learning)
- [[Doyen Sahoo]] — Salesforce AI 研究者(person / machine-learning)
- [[Steven C. H. Hoi]] — Salesforce AI VP of Research(person / machine-learning)
- [[Salesforce AI]] — Salesforce の AI 研究部門(organization)
- [[ARQ Foundation]] — ヨーロッパの AI 政策・安全保障シンクタンク。Europe 2031 の主執筆機関(organization / AI政策)
- [[ASML]] — オランダの EUV リソグラフィ装置メーカー。世界唯一の EUV 装置サプライヤー。AI 地政学のキーノード(organization / 半導体 / 地政学)
- [[Sakana AI]] — 東京拠点の AI 研究企業。自然界アルゴリズムインスパイアの AI 研究。Conductor 論文の所属機関(organization / ai)
- [[Stefan Nielsen]] — [[Sakana AI]] 研究者。RL Conductor の共同筆頭著者(person / ai)
- [[Edoardo Cetin]] — [[Sakana AI]] 研究者。RL Conductor のコアアルゴリズム実装・共同筆頭著者(person / ai)
- [[Yujin Tang]] — [[Sakana AI]] 研究者。Conductor プロジェクトリーダー(person / ai)
- [[Tingzhu Bi]] — [[Peking University]] 研究者。JustDiag 第一著者。AIOps・クラウドデータセンター障害診断(person / aiops)
- [[Xinrui Jiang]] — [[Peking University]] 研究者。JustDiag 共著者(person / aiops)
- [[Xun Zhang]] (PKU) — [[Peking University]] 研究者。JustDiag 共著者(person / aiops)
- [[Pengcheng Su]] — [[Peking University]] 研究者。JustDiag 共著者(person / aiops)
- [[Congjie He]] — [[University of Edinburgh]] 研究者。JustDiag 共著者(person / aiops)
- [[Jinglin Li]] — [[Beijing University of Posts and Telecommunications]] 研究者。JustDiag 共著者(person / aiops)
- [[Meng Ma]] — [[Peking University]] 研究者。[[Ping Wang]] とともに AIOps グループを形成。JustDiag 対応著者(person / aiops)
- [[Beijing University of Posts and Telecommunications]] — 北京郵電大学。情報通信・CS 分野の中国重点大学(organization)
#### SRE NEXT 2023 Runbook スライド (2026-06-23)
- [[Robert Treat]] — [[OmniTI]] CEO。SREcon16 で Less Alarming Alerts! を発表し、アラート追加前にビジネス影響・修復手順・通知先・予防可能性を問う実践を提示(person / sre / alert-management)
- [[OmniTI]] — Web アプリケーションとインフラストラクチャの大規模運用を支援する技術サービス企業。Less Alarming Alerts! の発表背景となる 24x7 運用組織(organization / sre)
- [[Sohei Iwahori]] — [[GREE, Inc]] のインフラ / Monitoring Unit Leader。SRE NEXT 2023 で Runbook とアラート振り分けの実装事例を発表(person / sre)
- [[GREE, Inc]] — ゲーム系ワークロードを中心にインフラ運用を行う企業。Runbook 整備とアラート追加ガイドラインの事例組織(organization / sre)
- [[Wei Zhang (Beihang)]] — Beihang University 所属、mABC の第一著者(EMNLP Findings 2024)
- [[Hongcheng Guo]] — Beihang University 所属、mABC 責任著者(EMNLP Findings 2024)
- [[Cloudwise]] — 中国 AIOps 企業、mABC 研究の産業パートナー
- [[Moshe Zadka]] — SRE 実践者・Python コントリビュータ。SREcon22 Americas でアラート品質のコストモデルを発表(person / sre / alert-management)
- [[Paige Cruz]] — Chronosphere シニアデベロッパーアドボケイト、元 SRE(New Relic / InVision / Lightstep / Weedmaps)。SREcon23 Americas で Alert Triage Hour of Power を発表(person / sre / observability)
- [[Chronosphere]] — クラウドネイティブオブザーバビリティプラットフォーム企業(organization / observability)
- [[Kristin Smith]] — Campspot DevOps Services チームリード。SREcon22 Americas でアラートポリューションと光害のアナロジーを発表(person / sre / observability)
- [[Campspot]] — アウトドアホスピタリティ向けソフトウェアスタートアップ(organization / outdoor-hospitality)
- [[Ren Xinchi]] — Alibaba Group GOC シニアエンジニア、ビジネスモニタリングの設計・運用担当(person / sre / monitoring)
- [[Hammurabi]] — Alibaba のビジネスモニタリング用 CMDB。ビジネス機能と P1〜P4 優先度・担当者をマッピング(product / monitoring / cmdb)
- [[Matt Bostock]] — Cloudflare プラットフォームオペレーションエンジニア。ロンドン Prometheus ミートアップ主催者(person / sre / monitoring)
- [[Dong Wang]] — Baidu プリンシパルアーキテクト。SRE チームを率い異常検知・障害自動修復に従事(person / sre / anomaly-detection)
- [[Baidu]] — 中国最大の検索エンジン企業。10 億ユーザー超(organization / internet / china)
#### OncallX (ASE 2025) (2026-06-24)
- [[Ruowei Fu]] — [[Nankai University]] 所属研究者、OncallX 筆頭著者(person / aiops / on-call)
- [[OncallX]] — Nankai University / ByteDance 共同開発のオンコール自動化システム。LLM × マルチエージェント協調でインシデント対応とチケットトリアージをカバー(product / aiops / on-call)
#### VCCL (arXiv 2026) (2026-06-26)
- [[Mingjun Zhang]] — [[Infrawaves]] 所属、VCCL 論文の共同筆頭著者(Xiaohe Hu と並列)(person / distributed / gpu-training)
- [[Infrawaves]] — (更新) VCCL の設計・24K GPU 本番展開組織として詳細を追記。訓練スループット最大 5.28% 向上・GPU 待機時間約 90% 削減(organization / distributed / gpu-training)
- [[Menghao Zhang]] — (更新) VCCL 論文の対応著者として追記。[[Beihang University]] 所属(person / distributed / rdma)
#### SREはサイバネティクスの夢をみるか (IOTS2025) (2026-06-26)
- [[坪内佑樹]] — さくらインターネット研究所研究員、京都大学博士(情報学)。テレメトリスケーリング・SRE の工学的体系化を研究(person / sre / telemetry)
- [[さくらインターネット研究所]] — さくらインターネット株式会社の研究開発部門。クラウドインフラ・テレメトリー・SRE 研究(organization / cloud)
#### Symptom-based Alerting for ML (SREcon23 EMEA) (2026-06-25)
- [[Lina Weichbrodt]] — ML フリーランス・コンサルタント、元 Zalando シニアリサーチエンジニア。30 以上の ML モデルを本番運用(person / ml-engineering)
#### デバッギング・性能解析・フィードバック 6 論文 (2026-06-26)
- [[Andreas Zeller]] — Universität des Saarlandes 計算機科学教授。デルタデバッギング(ddmin/dd)の提案者(person / debugging / testing)
- [[Ghassan Misherghi]] — UC Davis。HDD(階層的デルタデバッギング)筆頭著者(person / debugging)
- [[Zhendong Su]] — UC Davis 教授。HDD 共著者、プログラミング言語・ソフトウェアテスト(person / testing / programming-languages)
- [[Michael Chow]] — University of Michigan / Facebook。Mystery Machine 筆頭著者(person / performance-analysis)
- [[David Meisner]] — Facebook。Mystery Machine 共著者、エネルギー効率的データセンター研究(person / systems)
- [[Jason Flinn]] — University of Michigan 教授。分散システム・オペレーティングシステム(person / distributed-systems)
- [[Thomas F. Wenisch]] — University of Michigan 教授。コンピュータアーキテクチャ・性能解析(person / computer-architecture)
- [[George Candea]] — EPFL 教授。Failure Sketching(Gist)共著者、信頼性・自動回復(person / reliability / debugging)
- [[Jürgen Cito]] — TU Wien 助教授(元 University of Zurich)。FDD 論文筆頭著者、ソフトウェア性能工学(person / software-engineering / devops)
- [[Philipp Leitner]] — University of Gothenburg 教授。FDD 共著者、クラウド性能(person / software-engineering / cloud)
- [[Harald C. Gall]] — University of Zurich 教授。FDD 共著者、ソフトウェア進化(person / software-engineering)
- [[Colin Scott]] — UC Berkeley。DEMi 筆頭著者、分散システムデバッギング(person / distributed-systems / debugging)
- [[Scott Shenker]] — UC Berkeley / ICSI 教授。DEMi 共著者、ネットワーク・分散システム(person / networking / distributed-systems)
- [[George Necula]] — UC Berkeley 教授。DEMi 共著者、プログラム検証(person / programming-languages)
- [[Gist]] — EPFL 開発の障害スケッチングツール。協調解析と HW ウォッチポイントで本番障害の根本原因を 2.4% オーバーヘッドで近似診断(product / debugging / root-cause-analysis)
#### AIOps RCA/FI/OpsQA 7 papers batch ingest (2026-06-27)
- [[SparseRCA]] — テスト環境の疎トレースに対応した教師なし RCA。ExL パターンベース分解 + パーソナライズド PageRank(product / aiops / rca)
- [[FaultInsight]] — Microsoft のハイパースケールデータセンターホスト障害解釈可能診断フレームワーク(product / aiops / fault-diagnosis)
- [[LoFI]] — ログからの障害指示情報抽出。スパン予測モデル(product / aiops / log-analysis)
- [[iKnow]] — 意図誘導型クラウド運用 RAG チャットボット(product / aiops / opsqa)
- [[CausalBench]] — 介入的因果学習の障害箇所特定ベンチマーク(product / aiops / benchmark)
- [[ResilienceGuardian]] — マイクロサービスの障害耐性劣化変更検知(product / aiops / resilience)
- [[Xinrui Jiang]] — 清華大学。G-Cause 筆頭著者(person / aiops)
- [[Meng Ma]] — G-Cause 共著者(person / aiops)
- [[Ping Wang]] — G-Cause 共著者(person / aiops)
- [[Tingzhu Bi]] — Microsoft。FaultInsight 筆頭著者(person / aiops)
- [[Junjie Huang]] — FaultInsight 共著者(person / aiops)
- [[Saurabh Jha]] — 介入的因果学習 RCA 共著者(person / aiops)
- [[Saurabh Bagchi]] — Purdue University。介入的因果学習 RCA、CausalBench(person / aiops)
- [[Guanglei He]] — ResilienceGuardian 共著者(person / aiops)
- [[Zhihan Jiang]] — CUHK。LoFI 筆頭著者(person / aiops / log-analysis)
- [[Guangba Yu]] — SYSU。iKnow 筆頭著者(person / aiops / opsqa)
- [[Zhenhe Yao]] — 清華大学。SparseRCA 筆頭著者(person / aiops / rca)
#### The Morning Paper on Operability (blog.acolyer 2016) (2026-06-27)
- [[Adrian Colyer]] — Accel ベンチャーパートナー、元 Pivotal/VMware/SpringSource CTO。The Morning Paper 運営、CS 論文 400+ レビュー(person / venture-capital / research-communication)
- [[Roberto Natella]] — University of Naples Federico II 准教授。ソフトウェア障害注入の専門家。ThorFI 共著。AI システム FA/FI サーベイ(TOSEM 2025)共著(person / fault-injection / software-testing)
#### マイクロサービス RCA/FL 10 論文一括 ingest (2026-06-27)
- [[Zhiqiang Xie]] — Stanford University。Cloud Atlas 筆頭著者(person / fault-localization / llm)
- [[Yujia Zheng]] — Carnegie Mellon University。Cloud Atlas 共著者(person / causal-inference)
- [[Lizi Ottens]] — Stanford University。Cloud Atlas 共著者(person)
- [[Wenxiao Chen]] — 清華大学。Chain-of-Event 共著者(person / aiops)
- [[Huai Jiang]] — eBay Inc.。Chain-of-Event 共著者(person / industry)
- [[Liangfei Su]] — eBay Inc.。Chain-of-Event 共著者(person / industry)
- [[GrayScope]] — サーバー OS のグレー障害箇所特定システム。Shenglin Zhang ほか(product / fault-localization)
- [[Di Weng]] — Zhejiang University。RCInvestigator 共著者(person / visualization)
- [[Yingcai Wu]] — Zhejiang University。RCInvestigator 共著者(person / visualization)
- [[CSIRO Data61]] — CSIRO の研究機関。MicroIRC 共著者所属(organization / research)
- [[Tanner Lund]] — Microsoft Azure PRSE。SREcon19 APAC「A Tale of Two Postmortems」登壇者。Human Factors/Resilience Engineering 視点からポストモーテム実践を批判的再設計した(person / sre / microsoft)
- [[Sue Lueder]] — Google SRE Program Manager(2014 年入社)。インシデント分析プログラム設計者。SREcon 2015 登壇(person / sre / google)
- [[Nathan Hoffman]] — Etsy Software Engineer(2016年時点)。Etsy ポストモーテムファシリテーション教育修了者。SREcon16 Europe「Accident Models in Post Mortems」共同登壇(person / sre / etsy / postmortem)
- [[Miriam Lautner]] — Etsy Engineer(2016年時点)、Recurse Center 2014 卒。SREcon16 Europe「Accident Models in Post Mortems」で事故モデル理論パートを担当(person / sre / etsy / postmortem)
#### Running Excellent Retrospectives (SREcon19 Americas) (2026-06-28)
- [[Courtney Eckhardt]] — Heroku(Salesforce)SRE、@hashoctothorpe(she/her)。ポストモーテムファシリテーション言語の体系化。Miller's Law・contributing factor discovery・Why/You→How/What 変換(person / sre / heroku / facilitation)
- [[Lex Neva]] — Fastly SRE(he/him)。SRE Weekly 運営者(@SREWeekly)。Courtney Eckhardt と SREcon19 Americas チュートリアルを共同担当(person / sre / fastly)
- [[Fastly]] — エッジクラウドプラットフォーム(CDN)企業。Lex Neva の所属組織(organization / cdn)
#### Retrospectives for Humans (SREcon19 APAC) (2026-06-28)
- [[Courtney Eckhardt]] — (see above)
- [[Heroku]] — Salesforce 傘下のクラウドプラットフォーム(PaaS)。contributing factor discovery の実践組織(organization / cloud-platform)
#### Turning an Incident Report into a Design Issue with TLA+ (SREcon23 Americas) (2026-06-28)
- [[Finn Hackett]] — UBC PhD 学生。プログラミング言語・フォーマル検証研究。TLA+ インシデント分析ワークフローの提案者(person / formal-verification)
- [[Markus A. Kuppe]] — Microsoft Research プリンシパルリサーチエンジニア。TLA+ プロジェクト 10 年以上のエンジニア(person / formal-verification)
- [[Joshua Rowe]] — Microsoft プリンシパルエンジニア。Azure CosmosDB ドメインエキスパート、TLA+ モデル構築者(person / distributed-systems)
- [[Azure CosmosDB]] — Microsoft Azure のプラネットスケール KV ストア。5 段階整合性レベルを持ち TLA+ モデルが公開されている(product / database / distributed-systems)
#### Far from the Shallows (SREcon23 Americas) (2026-06-28)
- [[Courtney Nash]] — Verica Head of Research。Internet Incident Librarian。The Void(1 万件超の公開インシデントレポート DB)主宰。Duration/Severity/RCA 批判、インシデントストーリー提唱(person / sre / incident-research)
- [[Verica]] — カオスエンジニアリング・インシデント研究プラットフォーム企業。The Void データベースを通じた産業横断研究を推進(organization / sre / chaos-engineering)
#### インシデントキーメトリクスによるインシデント対応の改善 (SRE Kaigi 2025) (2026-06-28)
- [[Narimichi Takamura]] — [[Topotal]] CEO / SRE。@nari_ex。SRE Kaigi 2025 で MTTR の統計的限界と TTX メトリクスの体系的定義を発表(person / sre / topotal)
#### Human Factors in the Age of AI Ops (SREcon26 Americas) (2026-06-28)
- [[Eddie Redick]] — [[CTC Ops]] 所属。SREcon26 Americas で "Commanding the Chaos" フレームワーク・Trust Spectrum・Trust Triangle を提唱。"You don't rise to the level of your architecture; you fall to the level of your systems thinking."(person / sre / aiops / human-factors)
- [[CTC Ops]] — [[Eddie Redick]] の所属組織。SREcon26 Americas(2026-03-25)発表元(organization / sre / aiops)
#### Postmortem as a textbook (SpeakerDeck, 2023-02-09) (2026-06-28)
- [[KATO Toshiya]](新規) — LINE株式会社 Embedded SRE。@maruloop。SRE主導のポストモーテム執筆会議手法を考案・実践。(person / sre / postmortem)
- [[LINE株式会社]](新規) — 日本のIT企業。LINEアプリ・スタンプ・サービス群を開発・運営。Embedded SRE体制。2023年10月にLYコーポレーションへ移行。(organization / sre)
#### Incident Metrics in SRE (O'Reilly, 2021) (2026-06-28)
- [[Štěpán Davidovič]](新規) — Google SRE。内部インフラ自動モニタリング担当。チェコ工科大学(プラハ)2010 年卒。O'Reilly レポート「Incident Metrics in SRE」著者。(person / sre / google / metrics)
#### 縮約,網羅,減算:科学者の仕事とは何か (認知科学 2021) (2026-06-28)
- [[岡ノ谷 一夫]](新規) — 東京大学所属の行動神経科学者・認知科学者。小鳥のさえずり・ヒトの音楽知覚・メタ認知を研究。縮約・網羅・減算の三項対立で現代の科学方法論を論じる。(person / cognitive-science)
- [[東京大学]](新規) — 日本の国立大学・研究機関。(organization)
#### Data Center Networking 基盤論文 5 本 (2026-06-29)
- [[Mohammad Al-Fares]](新規) — UCSD。Fat-Tree・Hedera の第一著者。データセンターネットワークトポロジと動的フロースケジューリングの研究。(person / networking)
- [[Amin Vahdat]](新規) — UCSD→Google VP of Engineering。Fat-Tree・PortLand・Hedera の共著者。データセンターネットワーク研究の中心人物。(person / networking / google)
- [[Albert Greenberg]](新規) — Microsoft Research→Azure Networking VP。VL2 の第一著者、DCTCP 共著者。(person / networking / microsoft)
- [[Mohammad Alizadeh]](新規) — Stanford→MIT CSAIL。DCTCP の第一著者。データセンター輻輳制御の先駆的研究。(person / networking / congestion-control)
- [[Radhika Niranjan Mysore]](新規) — UCSD。PortLand の第一著者。L2 データセンターファブリック設計。(person / networking)
- [[Barath Raghavan]](新規) — Williams College→USC。Hedera 共著者。(person / networking)
- [[Sivasankar Radhakrishnan]](新規) — UCSD。Hedera 共著者。(person / networking)
- [[VL2]](新規) — Microsoft Research のデータセンターネットワークシステム。Clos + VLB + ディレクトリサービス。(product / networking)
- [[James Hamilton]](更新) — VL2 共著者として追記。(person / microsoft / amazon)
#### Spanner: Google's Globally Distributed Database (OSDI 2012 / TOCS 2013) (2026-06-28)
- [[James C. Corbett]](新規) — Google エンジニア。Spanner の第一著者。外部一貫性のあるグローバル分散データベースの設計・実装を主導。(person / distributed / systems)
- [[Jeffrey Dean]](更新) — Spanner の共著者として貢献を追記。(person / google)
- [[Sanjay Ghemawat]](更新) — Spanner の共著者として貢献を追記。(person / google)
#### Memory in the Age of AI Agents (arXiv 2025) (2026-06-29)
- [[Yuyang Hu]](新規) — NUS 博士課程。エージェントメモリサーベイの筆頭著者。(person / NUS / agent-memory)
- [[MemGPT]](新規) — OS ページング機構に着想を得た階層的メモリ管理フレームワーク。Packer et al., 2023。(product / agent-memory / framework)
- [[Mem0]](新規) — グラフ+ベクトルストアのハイブリッドメモリフレームワーク。Chhikara et al., 2025。(product / agent-memory / framework)
- [[Arnaud Lawson]](新規) — Squarespace Senior SRE。Ceph Object Storage 本番化と SLO 実装を主導。SREcon19 Americas 登壇。(person / sre / squarespace)
- [[Squarespace]](新規) — ニューヨーク拠点のウェブサイト構築・ホスティング企業。SLO 実装事例を SREcon19 Americas で公開。ELK スタック変革実録を SREcon19 EMEA で公開。(organization / tech)
- [[Alex Hidalgo]](新規) — Squarespace Observability SRE。SREcon19 EMEA で ELK 変革実録を発表。SLO/エラーバジェット実践者。(person / sre)
- [[Alex Lee]](新規) — Squarespace Service Reliability SRE。SREcon19 EMEA 共同発表者。(person / sre)
- [[National University of Singapore]](更新) — エージェントメモリサーベイの筆頭著者機関として追記。(organization)
#### Project Silica: Towards Sustainable Cloud Archival Storage in Glass (SOSP 2023) (2026-06-29)
- [[Antony Rowstron]](新規) — Microsoft Research 上級研究主任。Project Silica 主要設計者。分散ストレージ・P2P システム(Pastry / PAST / Farsite)でも著名。(person / microsoft / storage / distributed)
- [[Project Silica]](新規) — Microsoft が開発するガラス媒体(溶融石英)クラウドアーカイバルストレージシステム。SOSP 2023 で初論文発表。aka.ms/Silica。(product / microsoft / storage)
#### Quantifying Empathy Through Service Level Objectives (SREcon18 Asia/Pacific, 2018) (2026-06-29)
- [[Ketan Gangatirkar]](新規) — [[Indeed]] VP of Engineering(Job Seeker)。SREcon18 Asia で SLO 設計における「共感の数値化」フレームワークを発表。9 年以上にわたって Indeed の求職者プロダクトを担当。(person / sre / slo)
- [[Indeed]](新規) — 求人情報アグリゲーター企業。「One search. All jobs.」を標語に複数求人サイトをクローリング。SLO 違反(Vietnam 向け 13 分停止)の実体験から SLO 設計を強化。(organization / sre)
#### Principled Performance Analytics (SREcon22 Americas) (2026-06-30)
- [[Brent Bryan]](新規) — [[Google]] Cloud SRE。[[Narayan Desai]] と共著で [[2σ手法]](ワークロードコホート+z スコアによる性能定常性検定)を SREcon22 Americas(2022-03-16)で発表。GCP Data Analytics の本番適用を担当。(person / sre / google / performance)
- [[Narayan Desai]](更新) — SREcon22 Americas 登壇・2σ手法実装・「SLO は実現不可能」の根本批判を追記。
#### The Map Is Not the Territory: How SLOs Lead Us Astray (SREcon19 EMEA) (2026-06-30)
- [[Narayan Desai]](新規) — [[Google]] SRE(別名「Nora」, @nldesai)。SLO の 4 ユースケース分類・テール管理への SLO 不適用論・SLO Algebra の未解決問題を提示。SREcon19 EMEA (2019-10-03) 登壇。(person / sre / google / slo)
#### Not All Minutes Are Equal (SREcon23 Americas) (2026-06-30)
- [[Michael Goins]](新規) — [[Capital One]] SRE 組織変革エンジニア。SREcon23 Americas 登壇。(person / sre / slo)
- [[Troy Koss]](新規) — [[Capital One]] 企業 SRE 戦略リード。SREcon23 Americas 登壇。(person / sre / slo)
- [[Capital One]](更新) — LLM 推論研究に加え SRE 実践でも登場。SLO 採用失敗の構造分析と改善戦略の実践拠点。(organization / sre / fintech)
#### HPC Downtime Budgets (SREcon16 Europe) (2026-06-30)
- [[Cory Lueninghoener]](新規) — [[Los Alamos National Laboratory]] HPC 設計グループリーダー。エラーバジェットを HPC に適応した「ダウンタイム予算」の考案者。SREcon16 Europe 登壇。(person / sre / hpc)
- [[Los Alamos National Laboratory]](新規) — 米国 DOE / NNSA の国立研究所。HPC 規模は約 36,000 ノード・120 万コア・110 PB Lustre。ダウンタイム予算の実施拠点。(organization / hpc / research)
#### Run, Walk, Crawl, or How We Failed Our Way to SLO Readiness (SREcon25 EMEA) (2026-06-30)
- [[Rob Durst]](新規) — [[Spring Health]] SRE。Salt Lake City, Utah 在住。SREcon25 EMEA で SLO 導入の 4 度の失敗と成功を発表、「信頼性イニシアチブ・フレームワーク」提唱。(person / sre / slo)
- [[Spring Health]](新規) — 米国メンタルヘルス・テクノロジー企業(ハイパーグロース・スタートアップ)。2022: 50 eng/0 SRE → 2025: 200 eng/8 SRE・3,000 万 req/day。エラーバジェット起点コードフリーズ RFC 承認済み。(organization / startup / healthcare)
#### AI Assistants for Incident Lifecycle SLR (arXiv 2024) (2026-06-30)
- [[Dahlia Ziqi Zhou]](新規) — [[York University]] EASE ラボ院生。[[Marios Fokaefs]] 指導下でマイクロサービス AI 支援インシデント管理を研究。(person / aiops / microservice)
- [[Marios Fokaefs]](新規) — [[York University]] EASE ラボ主宰教員。マイクロサービス・クラウド信頼性・AIOps 専門。(person / aiops / microservice)
- [[York University]](新規) — カナダ・オンタリオ州トロントの公立大学。EASE ラボ(Fokaefs 主宰)がマイクロサービス信頼性研究を実施。(organization / university / canada)
#### Measuring Availability the Player Focused Way (SREcon25 Americas) (2026-06-30)
- [[Maxfield Stewart]](新規) — [[Riot Games]] Technical Director: Live Operations。SREcon25 Americas で Player Minutes SLO と Player Journey フレームワークの導入事例を発表。2021 年の可用性文化変革を主導。(person / sre / gaming)
- [[Riot Games]](新規) — ゲーム企業(LoL, Valorant 等)。MAU 1 億超・DAU 3,000 万超・CCU 200 万超・800+ サービス・50+ シャード。2021-2024 年に Player Journey SLO 導入を実施し可用性 99% を達成。(organization / gaming / sre)
- [[Derek Defields]](新規) — [[Riot Games]] CTO。2021 年に Maxfield Stewart に 12 ヶ月での可用性アカウンタビリティ確立を命じた。(person / gaming / cto)
#### X-lifecycle Learning for Cloud Incident Management using LLMs (FSE 2024) (2026-06-30)
- [[Aditya Singh]](新規) — [[Microsoft]] 所属研究者(
[email protected])。FSE 2024 X-lifecycle Learning 論文の共著者。(person / aiops / microsoft)
#### Keys to SRE (SREcon14, 2014) (2026-07-01)
- [[Ben Treynor Sloss]](更新) — SREcon14 2014 講演を sources に追加。「13 のキー」公開と「ローンチオンブラック」ルール・飽き性エンジニアの自動化インセンティブ・Wheel of Misfortune・移植可能性の nuclear option を追記。status seed → developing に昇格。(person / sre)
#### Incident Management and Chatops @ Netflix Feat Scorebot (SREcon16, 2016) (2026-07-01)
- [[Al Tobey]](新規) — [[Netflix]] SRE。Scorebot(Go 製 ChatOps ボット)の開発者。SREcon16(2016-03)で Scorebot による Netflix のインシデント管理 ChatOps 実践を発表。(person / sre / netflix)
- [[Netflix]](更新) — SREcon16 Scorebot 発表・[[Al Tobey]] を wiki 内言及に追加。(organization)
#### Incident Response @ FB, Facebook's SEV Process (SREcon16 Europe, 2016) (2026-07-01)
- [[Gareth Eason]](新規) — [[Facebook]] プロダクションレビュー(EMEA)運営者。HEAnet → Google → Facebook のキャリア。SREcon16 Europe(2016-07)で SEV Process を発表。(person / sre / facebook)
- [[Facebook]](更新) — SREcon16 Europe の SEV Process/Production Review 節を追加。Discoverer=Owner・意図的過大分類・メトリクスゲーミング警告・canary インシデントを追記。(organization / sre)
- [[Jay Parikh]](更新) — Eason 講演での Production Review 出席言及を追記(「head of infrastructure」表記の差異を注記)。(person / sre / facebook)
#### Incident Response in Unfamiliar Sociotechnical Systems (SREcon20 Americas, 2020) (2026-07-01)
- [[Morgan Collins]](新規) — [[Salesforce]] Principal SRE(Incident Response and Analysis)。技術エンジニアリング・運用分野で約20年の経験。SREcon20 Americas(2020-12-07)で ICS の起源・民間企業への適応・Warm Blanket Fallacy を発表。(person / sre / incident-management)
- [[Salesforce]](新規) — CRM クラウドプラットフォーム企業。[[Heroku]] を傘下に持つ。SRE・インシデント対応の実務組織として登場(AI研究部門の [[Salesforce AI]] とは別法人格上の区分)。(organization / sre / cloud)
- [[Pedro Canahuati]](更新) — Eason 講演での Production Staff Review 出席言及を追記(「head of engineering」表記の差異を注記)。(person / sre / facebook)
#### When Systems Flatline—Enhancing Incident Response with Learnings from the Medical Field (SREcon21, 2021) (2026-07-01)
- [[Sarah Butt]](新規) — [[Salesforce]] SRE。SREcon21(2021-10-14)で医療分野(ACLS/ATLS/WHO 手術チェックリスト)の実践を SRE インシデント対応に応用する講演を発表。(person / sre)
- [[Salesforce]](更新) — Sarah Butt の SREcon21 講演を wiki 内言及に追加。(organization / sre / cloud)
#### Dashboards and Runbooks: Scrapbooking for Engineers (SREcon22 Asia/Pacific, 2022) (2026-07-01)
- [[Colin Douch]](新規) — [[Cloudflare]] Observability Platform Team Tech Lead。鉱業出身、Observability 分野で約10年の経験。USENIX SREcon22 Asia/Pacific(2022-12-07)でダッシュボード・ランブックの肥大化病理と改善策を発表。(person / sre / observability / cloudflare)
- [[Cloudflare]](更新) — Colin Douch の SREcon22 APAC 講演(ダッシュボード・ランブックのライフサイクル論)を追加。(organization)
#### Epic Incidents of History: The 1979 NORAD Nuclear Near Miss (SREcon23 Americas, 2023) (2026-07-01)
- [[Nick Travaglini]](新規) — [[Honeycomb.io]] Technical Customer Success Manager。USENIX SREcon23 Americas(2023-03-21)で1979年 NORAD 核近接ミス事件の歴史的分析を発表。(person / sre / observability / honeycomb)
- [[Honeycomb.io]](新規) — オブザーバビリティ製品を提供する企業。Nick Travaglini が在籍。(organization / sre / observability)
- [[Vannevar Bush]](更新) — 軍産学複合体("Iron Triangle")の主導者としての側面を追加。SAGE・NORAD へ至る軍事デジタル技術発展の土台を作った歴史的文脈を記述。(person / information-science)
#### Incident Commanders (SREcon23 Americas, 2023) (2026-07-01)
- [[Vanessa Huerta Granda]](更新) — SREcon23 Americas(2023年、[[Jeli]] 在籍時)に [[Emily Ruppe]] と共同発表した「Incident Commanders」を追加。IC/アナリスト役割区分・専任 IC チーム構築経験を追記。SREcon25/26(Enova 在籍時)の既存記述より時系列的に早い講演として整理。(person / sre / incident-management / incident-commander)
- [[Emily Ruppe]](更新) — SREcon23 Americas での [[Vanessa Huerta Granda]] との共同講演を追加。支援(support)畑からのキャリアパスを追記。(person / sre / incident-management)
- [[Jeli]](更新) — SREcon23 Americas 講演を sources/related に追加。Emily Ruppe・Vanessa Huerta Granda の在籍時期を整理。(organization / sre / incident-management)
#### If I Can Do It on an Ambulance, You Can Do It in an Office: Scalable Incident Response Using ICS (SREcon23 Americas, 2023) (2026-07-01)
- [[Thai Wood]](新規) — 独立コンサルタント、元 EMT(救急救命士)。[[Resilience Roundup]] 主宰。USENIX SREcon23 Americas(2023-03-23)で救急医療の ICS 経験をソフトウェアのインシデント対応に応用する講演を発表。(person / sre / incident-management)
- [[Resilience Roundup]](新規) — Thai Wood が主宰するレジリエンスエンジニアリング関連の記事サイト。(product / sre)
#### The World Blew Up But We're All Okay: Managing a massive-scale incident at Datadog (SREcon23 EMEA, 2023) (2026-07-01)
- [[Laura de Vesine]](更新) — SREcon23 EMEA での [[Laurent Bernaille]] との共同講演を追加。Slack/メールハンドル「silverrose」を追記。(person / sre / datadog)
- [[Laurent Bernaille]](新規) — [[Datadog]] のエンジニア。Kubernetes・コンテナネットワーキング・クラウドインフラ運用を専門とする。USENIX SREcon23 EMEA で [[Laura de Vesine]] と共同発表。(person / sre / datadog / kubernetes)
- [[Datadog]](更新) — 2023年3月8日の大規模マルチクラウドインシデントと「you build it, you run it」文化、IC ローテーション、インシデントアプリの節を追加(既存 AI Research 内容は温存)。(organization / sre / aiops)
- [[Kubernetes]](更新) — Datadog の親子クラスタ構成とその復旧順序に関する節を追加。(product / kubernetes)
#### The Incident Is The Way: Using Your Incidents to Win Reliability Investment (SREcon23 EMEA, 2023) (2026-07-01)
- [[Niall McCarthy]](新規) — [[Afterpay]] エンジニアリングリーダー、インシデント管理担当。USENIX SREcon23 EMEA(2023-10-11、ダブリン)で「The Incident Is The Way」を発表。(person / sre / incident-management / afterpay)
- [[Afterpay]](新規) — 後払い決済(Buy Now, Pay Later)サービスを提供する企業。Niall McCarthy が所属。(organization / fintech / sre)
#### Hard Choices, Tight Timelines: A Closer Look at Tradeoff Decisions during Incidents (SREcon24 Americas, 2024) (2026-07-01)
- [[Laura Maguire]](更新) — Trace Cognitive Engineering/OSU 所属を反映し役割フィールドを更新。Skip-level トレードオフ意思決定研究(vignette 法)を追加。(person / sre / resilience-engineering / tradeoff)
- [[Courtney Nash]](更新) — The Void の限界(推論過程の欠落)とトレードオフ研究への展開を追加。status を seed から developing に更新。(person / sre / incident-management / tradeoff)
#### The Critical Resource Is You: Practical Destressing for On-Call Engineers (SREcon26 Americas, 2026) (2026-07-01)
- [[Beth Adele Long]](新規) — Continuous Re-integration 主宰 / Adaptive Capacity Labs Principal。SRE ウェルネス・インシデント対応コーチング専門。元 New Relic・Jeli・Gruntwork。(person / sre / stress-management / on-call)
- [[Continuous Re-integration]](新規) — Beth Adele Long が主宰する SRE ウェルネス / インシデント対応コーチング事業体。(organization / sre / coaching)
#### Your System Has Recovered from an Incident, but Have Your Developers? (SREcon18 Americas, 2018) (2026-07-01)
- [[Jaime Woo]](新規) — 元ジャーナリスト、元 Shopify テクノロジーコミュニケーション責任者。インシデント後の人的回復を SRE 文脈で提起。(person / sre / incident-management / human-factors)
#### Epistemology of Incident Management (SREcon26 Americas, 2026) (2026-07-01)
- [[Jack Kingsman]](新規) — [[Atlassian]] シニア SRE。認識論的インシデント管理フレームワーク(5フェーズ Incident Loop・証拠 2×2・探索 3 パターン・仮説 3 条件・テスト 6 基準)を SREcon26 Americas で発表。(person / sre / atlassian)
- [[Atlassian]](更新) — シニア SRE [[Jack Kingsman]] の SREcon26 Americas 発表を関連ソースに追加。(organization / sre / incident-management)
#### Retrieval as Reasoning authors (2026-07-02)
- [[Haoliang Ming]](新規) — WeChat/Tencent 研究者。LLM-Wiki の筆頭著者。Retrieval-as-Reasoning パラダイムを提唱。(person / nlp / tencent)
#### AI impact on science paper authors (2026-07-03)
- [[Qianyue Hao]](新規) — 清華大学 電子工学部 BNRist。AI と科学への社会的影響を計量書誌学的に分析。(person / tsinghua / scientometrics)
- [[Fengli Xu]](新規) — 清華大学 電子工学部 BNRist。計算社会科学・AI 影響分析。(person / tsinghua / scientometrics)
- [[Yong Li]](新規) — 清華大学 電子工学部教授 / 中関村アカデミー。(person / tsinghua / scientometrics)
- [[James Evans]](新規) — シカゴ大学 知識ラボ・社会学部教授 / サンタフェ研究所。科学の社会学・AI と知識生産を専門とする。(person / uchicago / sociology / science-of-science)
#### ML Fleet Efficiency paper authors (2026-07-02)
- [[Arissa Wongpanich]](新規) — [[Google]] ソフトウェアエンジニア。ML Productivity Goodput(MPG)フレームワークの筆頭著者。(person / google / systems-for-ml)
- [[Vijay Janapa Reddi]](新規) — Harvard 大学教授(Google 在籍中に ML Fleet Efficiency 研究)。MLPerf/MLCommons 設立者の一人。(person / harvard / systems-for-ml)
- [[Borg]](新規) — [[Google]] の大規模クラスタ管理・ジョブスケジューラー。Kubernetes の前身。MPG の SG 計算基盤。(product / google / scheduler)
- [[Google]](更新) — ML フリート効率(TPU + Borg)の研究成果・MPG フレームワークを関連情報に追加。(organization)
#### Vedrfolnir 著者 (2026-07-06)
- [[Yuxuan Chen]](新規) — [[Beihang University]] 所属。Vedrfolnir の第一著者。(person / networking)
- [[Xiheng Li]](新規) — [[Beihang University]] 所属。Vedrfolnir の共著者。(person / networking)
- [[Fangzheng Jiao]](新規) — [[Beihang University]] 所属。Vedrfolnir の共著者。(person / networking)
- [[Chunming Hu]](新規) — [[Beihang University]] 所属。Vedrfolnir の共著者・上席研究者。(person / networking)
- [[Menghao Zhang]](更新) — Vedrfolnir の対応著者(従来 VCCL・Hawkeye の共著者として記録済み)。(person / networking)
- [[Hawkeye]](更新) — Vedrfolnir がネットワーク側分析に採用するテレメトリ収集・プロベナンス解析基盤。(product / networking)
#### INTFusion (IFIP Networking 2026) の著者・組織 (2026-07-06)
- [[Leonardo Alberro]](新規) — [[Universidad de la República]] 所属。INTFusion の筆頭著者。(person / networking)
- [[Matias Richart]](新規) — [[Universidad de la República]] 所属。INTFusion の共著者。(person / networking)
- [[Eduardo Grampin]](新規) — [[Universidad de la República]] 所属。INTFusion の共著者。(person / networking)
- [[Universidad de la República]](新規) — ウルグアイ、モンテビデオの国立大学。INTFusion 全著者の所属機関。(organization / education / uruguay)
#### IPDPS 2026 AI Accelerator 比較論文の著者・組織 (2026-07-06)
- [[Giacomo Brunetta]](新規) — [[University of Illinois Chicago]] + [[Argonne National Laboratory]] 所属。IPDPS 2026 AI アクセラレータ比較論文の筆頭著者。(person / uic / argonne / hpc)
- [[Cerebras]](新規) — ウェーハスケールエンジン(WSE)を製造するデータフロー AI アクセラレータ企業。CS-3 は Llama 3.1 8B で 3,609 tok/s (A100 比 16.8×) を達成。(organization / hardware / aiinfra)
- [[SambaNova]](新規) — 再構成可能データフローアーキテクチャ(RDA)を採用した AI アクセラレータ企業。SN40L は HBM + 大容量 DDR のハイブリッドメモリ構成。(organization / hardware / aiinfra)
#### SREcon22 EMEA Oncall スライドの登壇者・組織 (2026-07-13)
- [[Dave O'Connor]](新規) — [[Twilio]] VP Engineering、元 Google SRE Director (2004–2021)。SREcon22 EMEA で toxic exceptionalism 批判と SRE 価値命題を論じた登壇者。(person / sre / twilio / google)
- [[Twilio]](新規) — クラウド通信 API プラットフォーム企業。[[Dave O'Connor]] が VP Engineering として在籍。(organization / cloud / communications)
#### AgentTether (arXiv 2026) 著者 (2026-07-13)
- [[Chenyu Zhao]](新規) — [[Nankai University]] 所属。AgentTether の筆頭著者。同著者グループの先行研究 PROBE(arXiv:2605.08717、未取り込み)の筆頭著者でもある。(person / nankai / agent-repair)
- [[Shenglin Zhang]](更新) — AgentTether の責任著者としての役割を追記。(person / nankai)
- [[Dan Pei]](更新) — AgentTether の共著者としての役割を追記。(person / tsinghua)
- [[Chetan Bansal]](更新) — AgentTether の共著者としての役割を追記。(person / microsoft)
- [[Saravan Rajmohan]](更新) — AgentTether の共著者としての役割を追記。(person / microsoft)
- [[Minghua Ma]](更新) — AgentTether の共著者としての役割を追記。(person / microsoft)
- [[Wenwei Gu]](更新) — AgentTether の著者所属(Nankai University)が既存の LLMPrism 記録(CUHK)と食い違うため contradiction callout を追加。(person / 要検証)
- [[Yongqian Sun]](更新) — AgentTether の共著者としての役割を追記。(person / nankai)
#### SOUPS 2025 セキュリティインシデント要約論文 著者 (2026-07-13)
- [[Diana Kramer]](新規) — [[Google]] 所属。筆頭著者。(person / security)
- [[Lambert Rosique]](新規) — [[DataPhant]] 所属。唯一の社外共著者。(person / security)
- [[Ajay Narotam]](新規) — [[Google]] 所属。(person / security)
- [[Elie Bursztein]](新規) — [[Google]] 所属。[[Sec-Gemini]] ブログ記事の共著者でもある。(person / security)
- [[Patrick Gage Kelley]](新規) — [[Google]] 所属。(person / security)
- [[Kurt Thomas]](新規) — [[Google]] 所属。(person / security)
- [[Allison Woodruff]](新規) — [[Google]] 所属。(person / security)
- [[Google]](更新) — セキュリティインシデント要約へのLLM統合(SOUPS 2025)を追記。
#### COMET / ISSRE 2024 インシデントトリアージ論文 著者・組織 (2026-07-13)
- [[Ze Li]](新規) — [[Microsoft]] 所属。COMET の共著者。(person / microsoft)
- [[Jianhui Li]](新規) — [[Chinese Academy of Sciences]] 所属。COMET の共著者。(person / cas)
- [[Chinese Academy of Sciences]](新規) — 中国の国立研究機関。COMET 論文の共著者所属機関(CNIC)。(organization / china)
- [[Zexin Wang]](更新) — COMET の筆頭著者としての役割を追記。既存の AgentOps サーベイ論文の著者と同一所属(CNIC CAS / UCAS)。(person / cas)
- [[Minghua Ma]](更新) — COMET の corresponding author としての役割を追記。(person / microsoft)
- [[Chetan Bansal]](更新) — COMET の共著者としての役割を追記。(person / microsoft)
- [[Qingwei Lin]](更新) — COMET の共著者としての役割を追記。(person / microsoft)
- [[Dongmei Zhang]](更新) — COMET の共著者としての役割を追記。(person / microsoft)
- [[Yu Kang]](更新) — COMET の共著者としての役割を追記。(person / microsoft)
- [[Chaoyun Zhang]](更新) — COMET の共著者としての役割を追記。(person / microsoft)
- [[Saravan Rajmohan]](更新) — COMET の共著者としての役割を追記。(person / microsoft)
- [[Murali Chintalapati]](更新) — COMET の共著者としての役割を追記。(person / microsoft)
- [[Changhua Pei]](更新) — COMET の共著者としての役割を追記。(person / cas)
- [[Gaogang Xie]](更新) — COMET の共著者としての役割を追記。(person / cas)
- [[Microsoft]](更新) — COMET のオンライン本番展開(精度30%改善・TTM35%短縮)を追記。(organization)
#### CoTriage / TOSEM投稿版 関連実体 (2026-07-13)
- [[Yang Zhang (ByteDance)]](新規)、[[Xin Wu (ByteDance)]](新規)、[[Feng Wang (ByteDance)]](新規)、[[Zeyu Che]](新規)、[[Xiaozhou Liu (ByteDance)]](新規) — CoTriage 共著者。(person / bytedance)
- [[Ruowei Fu]]・[[Yu Zhang (ByteDance)]]・[[ByteDance]]・[[Yongqian Sun]]・[[Nankai University]]・[[Wenwei Gu]]・[[Shenglin Zhang]](更新) — CoTriage の役割・所属を追記。
#### PROBE / Debugging the Debuggers 関連実体 (2026-07-13)
- [[Yihang Lin]](新規)、[[Zhimin Chen]](新規) — PROBE 共著者。(person)
- [[Chenyu Zhao]]・[[Shenglin Zhang]]・[[Wenwei Gu]]・[[Yongqian Sun]]・[[Dan Pei]]・[[Chetan Bansal]]・[[Saravan Rajmohan]]・[[Minghua Ma]]・[[AIOpsLab]](更新) — PROBE の共著者・評価基盤としての役割を追記。
#### Build-bench / Can Language Models Go Beyond Coding 関連実体 (2026-07-13)
- [[Build-bench]](新規、product)、[[Open Build Service]](新規、product)、[[Weilin Jin]](新規、person) — Build-bench ベンチマークの構成要素・共著者。
- [[Chenyu Zhao]]・[[Shenglin Zhang]]・[[Yongqian Sun]]・[[Dan Pei]]・[[Chaoyun Zhang]]・[[Qingwei Lin]]・[[Chetan Bansal]]・[[Saravan Rajmohan]]・[[Minghua Ma]]・[[Nankai University]]・[[Peking University]]・[[Tsinghua University]]・[[Microsoft]](更新) — Build-bench の共著者・所属を追記。
#### LagRCA 関連実体 (2026-07-13)
- [[Junhua Kuang]]・[[Yimeng Zhang]]・[[Jintao Feng]]・[[Jingyu Wang]]・[[Liping Zhang]](新規、person)、[[LagRCA]](新規、product) — LagRCA 共著者・手法名。
- [[Shenglin Zhang]]・[[Yongqian Sun]]・[[Dan Pei]]・[[Nankai University]]・[[Alibaba Group]]・[[Tsinghua University]]・[[Sibo Xia]]・[[Wenwei Gu]]・[[Wei Li]](更新) — LagRCA の共著者・所属を追記(Wei Li は同姓同名の可能性で contradiction 追加)。
#### InsightTriage / ASE'26投稿版 関連実体 (2026-07-13)
- [[Weiguo Li]](新規、person) — InsightTriage 共著者。
- [[Ruowei Fu]]・[[Shenglin Zhang]]・[[Wenwei Gu]]・[[Yongqian Sun]]・[[Dan Pei]]・[[Nankai University]](更新) — InsightTriage(Huawei/ICVドメイン)の役割・所属を追記。
#### Aloha / FSE Companion'26 関連実体 (2026-07-13、欠落補完)
- [[Yujia Wu]](新規、person)、[[Jinghuan Ren]](新規、person) — Aloha 第 2・第 3 著者([[Nankai University]])。source ページのリンク切れを解消。
- [[Shenglin Zhang]]・[[Yongqian Sun]]・[[Chaoyun Zhang]]・[[Liqun Li]]・[[Wenwei Gu]]・[[Qingwei Lin]]・[[Dongmei Zhang]]・[[Saravan Rajmohan]]・[[Chetan Bansal]]・[[Minghua Ma]]・[[Nankai University]]・[[Microsoft]](更新) — Aloha 論文への言及が欠けていたため共著者としての役割を追記。
#### FoundRoot / ICSE '26 関連実体 (2026-07-13)
- [[Yuzhuo Yang]](新規、person) — FoundRoot 著者([[Tsinghua University]])。
- [[Zhe Xie]]・[[Zeyan Li]]・[[Xiao He]]・[[Shenglin Zhang]]・[[Longlong Xu]]・[[Tieying Zhang]]・[[Jianjun Chen]]・[[Rui Shi]]・[[Dan Pei]]・[[Tsinghua University]]・[[ByteDance]]・[[Nankai University]](更新) — FoundRoot(構造化深層思考によるRCA基盤モデル)への言及を追記。
#### OScope / ICSE-SEIP '26 関連実体 (2026-07-13)
- [[OScope]](新規、product)、[[Yuxin Sun]]・[[Li Shi]]・[[Cheng Huang]]・[[Guodong Yang]]・[[Luping Wang]](新規、person) — Alibaba本番OS障害診断フレームワークOScopeとその共著者。
- [[Yongxin Zhao]]・[[Wenwei Gu]]・[[Yongqian Sun]]・[[Shenglin Zhang]]・[[Dan Pei]]・[[Liping Zhang]]・[[Nankai University]]・[[Alibaba Group]]・[[Tsinghua University]](更新) — OScope論文への言及を追記。
#### PerfScout / ICSE-SEIP '26 関連実体 (2026-07-13)
- [[Qingliang Zhang]]・[[Yimin Zuo]]・[[Bowen Deng]]・[[Xiao Xiong]]・[[Mengyao Li]]・[[Huandong Zhuang]]・[[Ruiyuan Wan]](新規、person) — PerfScout共著者。
- [[Yongqian Sun]]・[[Shenglin Zhang]]・[[Dan Pei]]・[[Xidao Wen]]・[[Nankai University]]・[[Huawei Cloud]]・[[BizSeer]]・[[Alban Siffer]]・[[Tsinghua University]]・[[Wenwei Gu]](更新) — PerfScout(適応的ワークロード生成)への言及を追記。
#### TADBench / IEEE TSC 2025 関連実体 (2026-07-13)
- [[Minyi Shao]]・[[Kaiwen Yang]]・[[Xingda Li]]・[[Dongbiao He]]・[[Yanbiao Li]](新規、person) — TADBench(トレース異常検知ベンチマーク)共著者。
- [[Yongqian Sun]]・[[Nankai University]]・[[Shenglin Zhang]]・[[Xiaohui Nie]]・[[Dan Pei]]・[[Changhua Pei]]・[[Bowen Hao]](更新) — TADBenchへの言及を追記(Xiaohui Nie はcorresponding author)。
#### LogSage / FCS 2025 関連実体 (2026-07-13)
- [[Tianyu Cui]](新規、person) — LogSage(カーネルパニックRCA)筆頭著者([[Nankai University]])。
- [[Shenglin Zhang]]・[[Yongqian Sun]]・[[Yicheng Sui]]・[[Zeyu Che]]・[[Nankai University]]・[[ByteDance]](更新) — LogSage論文への言及を追記。
#### RefinedEdge / IEEE TSC 2025 関連実体 (2026-07-13)
- [[RefinedEdge]](新規、product)、[[Jiacheng Zhang]]・[[Guohua Liu]]・[[Shiqi Chen]]・[[Yutong Chen]](新規、person) — エッジ・クラウド知識強化MTSADフレームワークRefinedEdgeとその共著者。
- [[Shenglin Zhang]]・[[Yongqian Sun]]・[[Dan Pei]]・[[Minghua Ma]]・[[Chenyu Zhao]]・[[Nankai University]]・[[Alibaba Cloud]](更新) — RefinedEdge論文への言及を追記。
#### A Survey of DevOps Concepts and Challenges 関連実体 (2026-07-14)
- [[Leonardo Leite]](新規、person、筆頭著者)、[[Carla Rocha]]・[[Fabio Kon]]・[[Paulo Meirelles]](新規、person) — DevOpsサーベイ(ACM CSUR 2019)の著者陣。
- [[University of São Paulo]]・[[University of Brasília]]・[[Federal University of São Paulo]](新規、organization) — 著者所属のブラジルの大学。
- [[Dejan Milojicic]]・[[Hewlett Packard Labs]](更新) — 2019年のDevOpsサーベイ論文への共著者としての言及を追記。
#### OpenRCA 2.0: From Outcome Labels to Causal Process Supervision 関連実体 (2026-07-14)
- [[Yifan Yang]]・[[Jin'ao Shang]]・[[Qisheng Lu]]・[[Rui Wang]]・[[Songhan Zhang]]・[[Yuzhong Zhang]]・[[Boxi Yu]](新規、person) — OpenRCA 2.0 論文の共著者。Jin'ao Shang は Pinjia He とともに corresponding author(所属メールドメインが他と異なる)。Boxi Yu も所属メールドメイン(lero.ie)が他と異なる。
- [[Aoyang Fang]]・[[Pinjia He]]・[[Junjielong Xu]]・[[The Chinese University of Hong Kong, Shenzhen]]・[[OpenRCA]](更新) — 後続研究 OpenRCA 2.0 への言及を追記。
#### The Anatomy of a Large-Scale Hypertextual Web Search Engine 関連実体 (2026-07-15)
- [[Sergey Brin]]・[[Lawrence Page]](新規、person) — Google 検索エンジンプロトタイプの共同開発者。Page は PageRank の考案者。
- [[Stanford University]]・[[Google]](更新) — Brin・Page が Stanford 在籍中に開発した検索エンジン創業論文への言及を追記。
#### Valet: Efficient Data Placement on Modern SSDs 関連実体 (2026-07-15)
- [[Devashish R. Purandare]]・[[Peter Alvaro]]・[[Avani Wildani]]・[[Darrell D. E. Long]]・[[Ethan L. Miller]](新規、person) — Valet 論文の著者陣。
- [[Valet]](新規、product)、[[MongoDB]]・[[CacheLib]](新規、product)、[[zenfs]]・[[f2fs]](新規、repository)、[[Pure Storage]](新規、organization) — Valet 本体、評価対象・比較対象システム、著者所属企業。
#### Recursive Self-Improvement (LessWrong, 2008) 関連実体 (2026-07-15)
- [[Eliezer Yudkowsky]](新規、person) — LessWrong創設者、AI安全性研究者。「AI go FOOM」論の提唱者。
- [[Robin Hanson]](新規、person) — 経済学者。テイクオフの「局所性」を巡り Yudkowsky と FOOM debate を行った論敵。
- [[I. J. Good]](新規、person) — 数学者。「知能爆発」概念(1965)の先行提唱者。
- [[UC Santa Cruz]]・[[Emory University]]・[[Cloudflare]]・[[RocksDB]](更新) — Valet 論文の著者所属先・ケーススタディ対象としての言及を追記。
#### Can Large Language Models Generate Observability-Aware Code? 関連実体 (2026-07-15)
- [[Yongliang Tao]]・[[Pengfei Gao]]・[[Zhiyu Fan]]・[[Jue Zhang]](新規、person) — 論文の共著者。Tao が筆頭著者(Chongqing University)、他3名は Microsoft 所属。
- [[Hongyu Zhang]]・[[Chongqing University]]・[[Minghua Ma]]・[[Qingwei Lin]]・[[Saravan Rajmohan]]・[[Si Qin]]・[[Liqun Li]]・[[Yu Kang]]・[[Microsoft]](更新) — 本論文への共著者としての言及を追記。
#### AI 2040: Plan A — The Deal 関連実体 (2026-07-16)
- [[AI Futures Project]](新規、organization) — AI存亡リスクの予測・政策提言を行う研究非営利団体。「AI 2027」「AI Futures Model」に続く3作目として「AI 2040」を発表。
- [[Daniel Kokotajlo]](新規、person) — [[AI Futures Project]]所属。本文中で唯一フルネームが判明する著者/所属者。AI存亡リスクのタイムライン予測・リード期間調査を行う。
#### A New Golden Age for Computer Architecture 関連実体 (2026-07-17)
- [[John L. Hennessy]](新規、person) — Stanford University 元学長、Alphabet Inc. 会長。RISC アーキテクチャ共同発明者、2017年 ACM Turing 賞受賞。
- [[RISC-V]](新規、product) — UC Berkeley 発のオープンソース ISA。モジュール式標準拡張(M/A/F/D/C)と RISC-V Foundation によるコミュニティ管理を特徴とする。
- [[David A. Patterson]](更新) — UC Berkeley Pardee 名誉教授・Google Distinguished Engineer という役職、Hennessy との共同 Turing 賞受賞、本論文での RISC-I 開発経緯を追記。
- [[Google]](更新) — TPU v1 の内部構成(Systolic Array・Matrix Multiply Unit)とドメイン固有アーキテクチャとしての性能・エネルギー効率評価を追記。
#### ContextPilot 関連実体 (2026-07-18)
- [[ContextPilot]](新規、product) — [[University of Edinburgh]] 開発のコンテキストレベル KV キャッシュ再利用システム。MLSys 2026 Oral、OSS 公開。
- [[University of Edinburgh]](更新) — ContextPilot の開発拠点であることを追記。
- [[LMCache]](更新) — ContextPilot が完全一致方式の代表ベースラインとして比較(MultihopRAG でヒット率4.6%)。
- [[CacheBlend]](更新) — ContextPilot が近似 KV マッチングの代表ベースラインとして比較。精度劣化の報告値が CacheBlend 原論文と食い違う点を contradiction として追記。
- [[Mem0]](更新) — ContextPilot が LoCoMo ベンチマーク上でエージェントメモリの整列によるTTFT短縮を評価。
#### The Too-Much-Talent Effect 関連実体 (2026-07-18)
- [[Roderick I. Swaab]]・[[Michael Schaerer]]・[[Eric M. Anicich]]・[[Richard Ronay]]・[[Adam D. Galinsky]](新規、person) — 論文の共著者。Swaab が筆頭著者([[INSEAD]])、Anicich・Galinsky は [[Columbia University]]、Ronay は [[Vrije Universiteit Amsterdam]](論文中表記 VU University Amsterdam)所属。
- [[INSEAD]](新規、organization) — フランス・フォンテーヌブローの国際経営大学院。Swaab・Schaerer の所属機関。
- [[Columbia University]]・[[Vrije Universiteit Amsterdam]]・[[Singapore Management University]](更新) — 本論文の著者所属機関としての言及を追記。
#### AI生成テキスト分類器 関連実体 (2026-07-20)
- [[lyc8503]](新規、person) — ブログ記事著者。TF-IDF+SVMベースのAI生成テキスト検知器を週末プロジェクトとして開発。
- [[AITextDetector]](新規、repository) — [[lyc8503]] が開発したAI生成テキスト検知器。GitHubリポジトリとブラウザ推論デモから成る。
#### Adversarial dynamical systems 関連実体 (2026-07-20)
- [[Matthew J. Colbrook]](新規、person) — [[University of Cambridge]] DAMTP。Koopman作用素のデータ駆動スペクトル計算(ResDMD・mpEDMD等)を専門とする応用数学者、本論文の責任著者。
- [[Igor Mezić]](新規、person) — [[UC Santa Barbara]]。Koopman作用素論の創始者の一人で応用Koopman理論の提唱者。
- [[Alexei Stepanenko]](新規、person) — [[University of Cambridge]]。初期段階の議論・遷移補題の予備検討に貢献。
- [[UC Santa Barbara]](更新) — [[Igor Mezić]] によるKoopman作用素論・力学系解析の研究拠点であることを追記。
- [[University of Cambridge]](更新) — lint-stub から実体化。DAMTPに[[Matthew J. Colbrook]]・[[Alexei Stepanenko]]が在籍することを追記。
### 2026-07-20 FailSafe ingest-paper
- [[Ziyi Xu]](新規、person) — [[Shanghai Jiao Tong University]]。FailSafe(耐障害 LLM サービング)筆頭著者。
- [[Zhiqiang Xie]](更新) — FailSafe(arXiv 2025、MLSys 2026 Oral)の共著者であることを追記。
- [[Swapnil Gandhi]](更新) — FailSafe の共著者であることを追記。訓練(ReCycle)とサービング(FailSafe)双方の耐障害設計に取り組む研究者として位置づけ。
- [[Christos Kozyrakis]](更新) — FailSafe の共著者であることを追記。NVIDIA Research 所属も併記。
- [[Stanford University]](更新) — FailSafe 研究への関与を追記。
- [[Shanghai Jiao Tong University]](更新) — [[Ziyi Xu]] の所属として FailSafe への関与を追記。
- [[ReCycle]](更新) — 姉妹システム FailSafe(サービング向け)との対比を追記。
### 2026-07-20 In-House LLM Serving at Netflix ingest
- [[Triton Inference Server]](新規、product) — [[NVIDIA]]。Netflix の Model Scoring Service(MSS)で GPU 推論の共有バックエンドとして運用。Python/vLLM バックエンドの選択とバージョン整合、OpenAI 互換フロントエンドを収録。
- [[Netflix]](更新) — 既存 JVM サービングシステム + MSS/Triton 上での LLM 内製サービング事例を追記。TensorRT-LLM→vLLM 移行を報告。
- [[vLLM]](更新) — Netflix の paved-path エンジン採用事例(運用適合性が判断基準)、V0→V1 移行による logits processor バッチレベル化を追記。
- [[TensorRT-LLM]](更新) — Netflix が旧 paved-path エンジンとして使用していたことと、vLLM への移行理由を追記。
- [[NVIDIA]](更新) — Triton Inference Server の本番運用事例(Netflix)を追記。
### 2026-07-20 Niyama ingest-paper
- [[Kanishk Goel]](新規、person) — [[Microsoft Research]] India。Niyama の筆頭著者。
- [[Jayashree Mohan]](新規、person) — Microsoft Research India。Niyama の共著者。
- [[Nipun Kwatra]](新規、person) — Microsoft Research India。Niyama の共著者。
- [[Ravi Shreyas Anupindi]](新規、person) — Microsoft Research India。Niyama の共著者。
- [[Ramachandran Ramjee]](新規、person) — Microsoft Research India。Niyama の共著者。Sarathi・Vidur・Etalon 等 LLM 推論サービング研究群の中心人物。
- [[Sarathi-Serve]](新規、product) — [[vLLM]] 上に構築された chunked-prefill スケジューラ。Niyama の拡張元として詳細を収録。
- [[Microsoft Research]](更新) — India ラボの LLM 推論サービング研究群(Niyama・Sarathi・Vidur・Etalon)と所属研究者を追記。
- [[vLLM]](更新) — Niyama による QoS 駆動スケジューリング拡張(Sarathi-Serve 経由)を追記。
### 2026-07-20 DuckDB ingest-paper
- [[Mark Raasveldt]]・[[Hannes Mühleisen]](新規、person) — [[CWI]] Database Architectures Group。DuckDBの共同開発者。前身[[MonetDBLite]]の開発者でもある。
- [[CWI]](新規、organization) — オランダ・アムステルダムの研究機関。[[DuckDB]]・[[MonetDBLite]]の開発拠点。
- [[DuckDB]](新規、product) — CWIが開発した組み込み型分析(OLAP)データベース。ベクトル化解釈実行・DataBlocksストレージ・シリアライザブルMVCCを収録。
- [[MonetDBLite]](新規、product) — DuckDBの前身にあたる、MonetDB派生の組み込み型分析データベース。
### 2026-07-20 DiDi #06(Query Execution Plans and Pipelining)ingest-slides
- [[Torsten Grust]](更新、person) — DiDi第6回「Query Execution Plans and Pipelining」を追加。パイプライン・シンク・パイプライン駆動ループの解説者として。
- [[DuckDB]](更新、product) — DiDi第6回の内容(実行プランのパイプライン分解・並列実行モデル)を講義教材節に追記。
- [[Universität Tübingen]](更新、organization) — DiDi第6回のソース・relatedを追加。
### 2026-07-20 DiDi #07(Vectorized Query Execution)ingest-slides
- [[Torsten Grust]](更新、person) — DiDi第7回「Vectorized Query Execution」を追加。DuckDB 1.4のC++ソースコード(ExpressionExecutor/VectorOperations/BinaryExecutor)を追跡する解説者として。
- [[DuckDB]](更新、product) — ベクトル・data chunk・morselの実行モデル、物理表現(FLAT/CONSTANT/DICTIONARY/SEQUENCE)、unified representation+C++テンプレートによるコード生成、SIMD化・分岐予測ペナルティを講義教材節に追記。
- [[Universität Tübingen]](更新、organization) — DiDi第7回のソース・relatedを追加。
### 2026-07-20 30分でわかるデータ指向アプリケーションデザイン(Data Engineering Study #18)ingest-slides
- [[Taro L. Saito]](新規、person) — 『データ指向アプリケーションデザイン』監訳者。Data Engineering Study #18 での登壇者として。
- [[Amazon Aurora (Database)]](更新) — 講演で紹介された classic Aurora(SIGMOD 2018)の gossip プロトコルによる2PC回避の平易な説明を[[分散トランザクション]]概念への参照として追記。
- [[DuckDB]](更新) — 列指向フォーマット(Parquet)のエコシステム対応例として言及されたことを追記。
### 2026-07-20 LLM hallucinations in the wild (arXiv 2605.07723) ingest-paper
- [[Zhenyue Zhao]]・[[Yihe Wang]]・[[Toby Stuart]]・[[Mathijs De Vaan]]・[[Paul Ginsparg]]・[[Yian Yin]](いずれも新規、person) — arXiv・bioRxiv・SSRN・PubMed Central のハルシネーション引用大規模監査研究の著者陣。Cornell University・Tsinghua University・University of California, Berkeley(Haas School of Business)に所属。
- [[Cornell University]]・[[University of California, Berkeley]]・[[Tsinghua University]](いずれも更新) — 本研究の著者所属先として related に追記。既存の AIOps/システム系ソースとは異なる science-of-science 分野の新規ソースを積み増した。
### 2026-07-20 Aurora DSQL: Scalable, Multi-Region OLTP (arXiv 2607.13276) ingest-paper
- [[Aurora DSQL]](新規、product) — disaggregated アーキテクチャを持つ AWS のサーバーレス・マルチリージョン分散 SQL データベース。[[Amazon Aurora (Database)]]とは別システムであることを明示。
- [[Marc Brooker]](更新、person) — Aurora DSQL の筆頭著者としての貢献を追記。
- [[Amazon Aurora (Database)]](更新) — Aurora DSQL との名称混同回避のための注記と、両システムのアーキテクチャ上の関係(ストレージ層を継承しない別設計)を追記。
### 2026-07-20 Using Lightweight Formal Methods to Validate a Key-Value Storage Node in Amazon S3 (SOSP 2021) ingest-paper
- [[ShardStore]](新規、product) — Amazon S3 のキーバリューストレージノード実装。LSM ツリー + エクステント外 shard データ配置 + soft updates 由来の Dependency 型によるクラッシュ整合性を持つ。
- [[James Bornholt]](新規、person) — 論文筆頭著者。Amazon Web Services 所属、Yggdrasil・Ferrite・Hyperkernel 等のストレージ/OS 形式検証研究の実績を持つ。
- [[Amazon Web Services]](更新) — ShardStore の開発・検証組織としての役割を追記。
### 2026-07-20 The Snowflake Elastic Data Warehouse (SIGMOD 2016) ingest-paper
- [[Snowflake Computing]](新規、organization) — マルチクラスタ・シェアードデータ・アーキテクチャを持つクラウドデータウェアハウス Snowflake の開発企業。
- [[Benoit Dageville]](新規、person) — 論文筆頭著者。
- [[Thierry Cruanes]](新規、person) — 共著者。Cloud Services層(クエリオプティマイザ・並行性制御・プルーニング)の設計に貢献。
- [[Marcin Zukowski]](新規、person) — 共著者。VectorWise/MonetDB/X100由来のベクトル化実行モデルの系譜を Snowflake に持ち込んだ人物と位置づけ。
- [[Amazon Web Services]](更新) — Snowflake の稼働基盤プラットフォームとしての役割、Aurora Limitless Databaseとの組織・アーキテクチャ思想の違いを追記。
### 2026-07-20 Dremel: Interactive Analysis of Web-Scale Datasets (VLDB 2010) ingest-paper
- [[Sergey Melnik]](新規、person) — 論文筆頭著者。
- [[Andrey Gubarev]](新規、person) — 共著者。
- [[Jing Jing Long]](新規、person) — 共著者。
- [[Geoffrey Romer]](新規、person) — 共著者。
- [[Shiva Shivakumar]](新規、person) — 共著者。
- [[Matt Tolton]](新規、person) — 共著者。
- [[Theo Vassilakis]](新規、person) — 共著者。
- [[MapReduce]](新規、product) — Dremelが補完対象とする既存のGoogleバッチ計算フレームワーク。並行して収録された [[@2004__OSDI__MapReduce - Simplified Data Processing on Large Clusters]] を実体ソースとして紐付け。
- [[Protocol Buffers]](新規、product) — Dremelのネストデータモデルの基盤となるGoogleのシリアライズ機構。
- [[Google]](更新) — 分散データベース基盤セクションにDremelのエントリを追記。
- [[Andrew Crotty]](新規、person) — Mach の共著者。[[Brown University]]・[[Carnegie Mellon University]] 二重所属。
- [[Mach]](新規、product) — 疎結合アーキテクチャによるメトリクス専用ストレージエンジン。
- [[Franco Solleza]](更新) — Mach の筆頭著者としてのエントリを追記。
- [[Nesime Tatbul]](更新) — Mach の共著者としてのエントリを追記。
- [[Stan Zdonik]](更新) — Mach の共著者としてのエントリを追記。
- [[Suman Karumuri]](更新) — Mach の共著者としてのエントリを追記。
- [[Brown University]](更新) — Mach の著者所属・評価データセット出所として追記。
- [[Carnegie Mellon University]](更新) — Andrew Crotty の所属として追記。
- [[Slack Technologies]](更新) — Mach 冒頭の障害事例(2020年5月12日アウテージ)・規模感の再引用として追記。
### 2026-07-21 Don't Predict, Prioritize: Rethinking GPU Reliability Assessment ingest-paper
- [[Difeng Ma]](新規、person) — HeaRank 論文の筆頭著者。CNIC/CAS・UCAS 所属、StepFun でリサーチインターンとして本研究を実施。
- [[Yuanwei Lu]](新規、person) — StepFun 所属の共著者。
- [[Quan Zhou]](新規、person) — CNIC/CAS 所属の共著者。
- [[Daxin Jiang]](新規、person) — StepFun 所属の共著者。
- [[Jingjing Li]](新規、person) — CNIC/CAS 所属の共著者。
- [[Changhua Pei]](更新) — HeaRank 論文の corresponding author としてのエントリを追記。
- [[Gaogang Xie]](更新) — HeaRank 論文の共著者としてのエントリを追記。
- [[Zexin Wang]](更新) — HeaRank 論文の共著者としてのエントリを追記。
- [[Yibo Zhu]](更新) — HeaRank 論文の共著者としてのエントリを追記。GPU クラスタスケジューリング・LLM推論に続く3つ目の研究軸。
- [[Dan Pei]](更新) — HeaRank 論文の共著者としてのエントリを追記。GPU ハードウェア信頼性という新ドメインへの拡張。
- [[Chinese Academy of Sciences]](更新) — HeaRank 論文の共著者所属機関としての言及を追記。
- [[University of Chinese Academy of Sciences]](更新) — HeaRank 論文の共著者所属機関としての言及を追記。
- [[Tsinghua University]](更新) — Dan Pei 経由の HeaRank 論文言及を追記。
- [[StepFun]](更新) — GPU クラスタ運用者・LLM 推論システム開発元に加え、GPU 信頼性研究への実データ提供元としての役割を追記。