# 論文キュレーション 2026-W33 前回実行(2026-W28)から 5 週間ぶりの巡回。会議側は SOSP・ISCA・KDD・ASE・ICWS・IWQoS・ICDCS・ICS・ISSTA・IMC の 10 会議回を新たに解析した。 ## 必読 スコア 80 以上。44 件。 ### [GALA: Graph-Augmented LLM Agents for Root Cause Analysis and Incident Response in Microservices](https://conf.researchr.org/track/ase-2026/ase-2026-research-track) - 会場/年: ASE 2026 / 2026 - スコア: 95 - 収集元: 会議 - 査読: 査読済み - 推薦理由: グラフ拡張した LLM エージェントによるマイクロサービスの根本原因分析とインシデント対応であり、コア関心の中心に位置する。profile が例外基準で名指しした GALA の査読版に相当し、arXiv 版から会議採択版への更新として読む価値が高い。障害依存グラフと LLM 推論の統合という方法論の好みにも合う。 - 関連: [[@2025__arXiv__GALA - Can Graph-Augmented Large Language Model Agentic Workflows Elevate Root Cause Analysis]] ・ [[@2024__FSE Companion__Exploring LLM-Based Agents for Root Cause Analysis]] ### [LMTracer: Fine-Grained and Real-Time Performance Profiling for Production LLM Systems](https://sigops.org/s/conferences/sosp/2026/accepted.html) - 会場/年: SOSP 2026 / 2026 - スコア: 95 - 収集元: 会議 - 査読: 査読済み - 推薦理由: 本番 LLM システムの細粒度かつリアルタイムな性能プロファイリングであり、AI インフラの可観測性というコア関心のど真ん中に位置する。本番環境を対象と明示している点は実運用規模での評価という好みに合致する。LLMPrism との比較軸としても押さえておきたい。 - 関連: [[@2025__DSN__LLMPrism - Black-box Performance Diagnosis for Production LLM Training Platforms]] ・ [[@2026__AI__Scalable and Energy-Efficient AI - System-Level Profiling of NVIDIA GPU Clusters for Distributed LLM Training]] ### [RCAgentBench: An Agent-Oriented Benchmark for Multimodal Root Cause Analysis in Microservices](https://iwqos2026.ieee-iwqos.org/detailed-technical-sessions) - 会場/年: IWQoS 2026 / 2026 - スコア: 93 - 収集元: 会議 - 査読: 査読済み - 推薦理由: マイクロサービスのマルチモーダル根本原因分析をエージェント指向で評価するベンチマークであり、AIOps 評価設計と多モーダル統合という好みの双方に合致する。最終回答のみでなくエージェントの行動を評価する設計であれば、推論プロセス志向という評価の癖に強く応える。RCAEval や Cloud-OpsBench と並べて読む価値が高い。 - 関連: [[@2025__WWW Companion__RCAEval - A Benchmark for Root Cause Analysis of Microservice Systems with Telemetry Data]] ・ [[@2026__arXiv__Cloud-OpsBench - A Reproducible Benchmark for Agentic Root Cause Analysis in Cloud Systems]] ### [StaR: Stateful Dynamic-Graph Root Cause Analysis through Memory-Enhanced Causality Discovery](https://www.semanticscholar.org/paper/c08684985da94b5c8448fa454c98140c87a8da4d) - 会場/年: KDD 2026 / 2026 - スコア: 92 - 収集元: メール+会議 - 査読: 査読済み - 推薦理由: マイクロサービスの多変量時系列に対し、動的トポロジと状態保持を組み込んだグレンジャー因果発見で根本原因を特定する。静的・記憶なしという既存の因果 RCA の前提を崩す点が新しく、コア関心の因果発見型 RCA そのものである。九つのデータセットと実世界データでの改善幅も具体的に示されている。 - 関連: [[@2025__arXiv__RADICE - Causal Graph Based Root Cause Analysis for System Performance Diagnostic]] ・ [[@2024__CCGRID__AlertRCA - Causality Enhanced Graph Representation Learning for Alert-Based Root Cause Analysis]] ### [Gleaner: A Semantically-Rich and Efficient Online Sampler for Microservice Diagnostics](https://conf.researchr.org/track/issta-2026/issta-2026-research-papers) - 会場/年: ISSTA 2026 / 2026 - スコア: 90 - 収集元: 会議 - 査読: 査読済み - 推薦理由: マイクロサービス診断のための意味情報に基づくオンライントレースサンプラであり、テレメトリ削減と障害診断を同時に扱うコア関心の直撃である。統計だけでなく意味・構造の事前知識で収集対象を絞る設計は方法論の好みにも合致する。オンライン動作を掲げる以上、収集オーバーヘッドの数値も期待できる。 - 関連: [[@2024__TSC__No More Data Silos - Unified Microservice Failure Diagnosis With Temporal Knowledge Graph]] ・ [[@2021__CloudIntelligence__MicroDiag - Fine-grained Performance Diagnosis for Microservice Systems]] ### [QProf: Fleetwide Transitive Cost Profiling of Warehouse-Scale Services](https://sigops.org/s/conferences/sosp/2026/accepted.html) - 会場/年: SOSP 2026 / 2026 - スコア: 90 - 収集元: 会議 - 査読: 査読済み - 推薦理由: 倉庫規模のサービス群に対しフリート全体で推移的なコストを帰属させる継続プロファイリング基盤であり、非侵入プロファイリングと実運用規模の評価という中核関心に真正面から合致する。呼び出し先まで遡ってコストを配賦する設計は、サービス間の責任配分という運用上の難問に答えるものである。今週の候補で最も優先度が高い。 - 関連: [[@2022__USENIX ATC__CRISP - Critical Path Analysis of Large-Scale Microservice Architectures]] ・ [[@2014__OSDI__lprof - A Non-intrusive Request Flow Profiler for Distributed Systems]] ### [QueueBreak: A Trace-to-Diagnosis Pipeline for Tail Latency in Agentic LLM Services](https://services.conferences.computer.org/2026/icws-program/) - 会場/年: ICWS 2026 / 2026 - スコア: 90 - 収集元: 会議 - 査読: 査読済み - 推薦理由: エージェント型 LLM サービスのテール遅延をトレースから診断まで一貫して扱うパイプラインであり、分散トレーシングと障害診断と AI インフラ可観測性が同時に交わる。待ち行列に着目した遅延要因の切り分けは、単なる異常検知でなく原因の説明に踏み込む設計であり評価の癖に合う。コア関心の中でも優先度が高い。 - 関連: [[@2026__arXiv__AgentTether - Graph-Guided Diagnosis and Runtime Intervention for Reliable LLM Agent Operations]] ・ [[@2019__WWW__ε-Diagnosis - Unsupervised and Real-time Diagnosis of Small-window Long-tail Latency in Large-scale Microservice Platforms]] ### [ActionNex: A Production-Grade System for Next-Best-Action Recommendation in Cloud Outage Management](https://conf.researchr.org/track/ase-2026/ase-2026-industry-showcase) - 会場/年: ASE 2026 / 2026 - スコア: 88 - 収集元: 会議 - 査読: 査読済み - 推薦理由: クラウド障害対応における次の一手を推薦する本番稼働システムであり、インシデント管理の自律化というコア関心に合致する。産業展示トラックであり実運用規模での運用知見が得られる可能性が高い。ベンダー自己申告値の扱いには注意が必要だが、実配備事例として価値がある。 - 関連: [[@2019__WWW__Outage Prediction and Diagnosis for Cloud Service Systems]] ・ [[@2026__UIUC PhD Thesis__Agentic Failure Management of Cloud Systems]] ### [BulletTime: Time Dilation for High-Fidelity Tracing](https://www.semanticscholar.org/paper/b1a5df7e0a4503209b89cb6a1bef6d106b179769) - 会場/年: ISCA 2026 / 2026 - スコア: 88 - 収集元: メール+会議 - 査読: 査読済み - 推薦理由: 時間を引き延ばすことで高忠実度トレーシングの観測オーバーヘッド問題に対処する提案であり、可観測性とテレメトリ工学というコア関心の中心にある。計装オーバーヘッドという代償を正面から扱う点も高く評価できる。 - 関連: [[@2025__ISSTA__Tracezip - Efficient Distributed Tracing via Trace Compression]] ・ [[@2025__ASPLOS__Mint - Cost-Efficient Tracing with All Requests Collection via Commonality and Variability Analysis]] ### [Enhancing Trace-Based Root Cause Analysis for Microservice Systems via Code Change Understanding](https://conf.researchr.org/track/ase-2026/ase-2026-research-track) - 会場/年: ASE 2026 / 2026 - スコア: 88 - 収集元: 会議 - 査読: 査読済み - 推薦理由: トレースにコード変更の理解を組み合わせた根本原因分析であり、コア関心の RCA と可観測性が重なる。変更起因障害という実運用で支配的な障害種別を扱い、意味情報で探索空間を絞る方針にも合致する。ChangeRCA との比較読みに向く。 - 関連: [[@2024__FSE__ChangeRCA - Finding Root Causes from Software Changes in Large Online Systems]] ・ [[@2024__ASE__Root Cause Analysis for Microservices based on Causal Inference - How Far Are We]] ### [HERO: Hypothesis-Centered Root-Cause Analysis for Microservice Incidents](https://conf.researchr.org/track/ase-2026/ase-2026-research-track) - 会場/年: ASE 2026 / 2026 - スコア: 88 - 収集元: 会議 - 査読: 査読済み - 推薦理由: 仮説を中心に据えたマイクロサービスインシデントの根本原因分析であり、最終回答だけでなく推論過程を構造化する点が評価の癖に強く合う。仮説生成と検証の分離は、箇所特定と原因同定を分けて評価したいという関心に直接応える。 - 関連: [[@2022__NeurIPS__Root Cause Analysis of Failures in Microservices through Causal Discovery]] ・ [[@2024__FSE__Chain-of-Event - Interpretable Root Cause Analysis for Microservices through Automatically Learning Weighted Event Causal Graph]] ### [On-site, Non-speculative Failure Diagnosis with CLODS](https://sigops.org/s/conferences/sosp/2026/accepted.html) - 会場/年: SOSP 2026 / 2026 - スコア: 88 - 収集元: 会議 - 査読: 査読済み - 推薦理由: 本番稼働中のシステムで推測に頼らない障害診断を行う。現地診断は計装オーバーヘッドと診断精度のトレードオフに正面から向き合う設計であり、根本原因分析というコア関心の中心にある。SOSP 採択という点でも信頼度が高い。 - 関連: [[@2015__SOSP__Failure Sketching - A Technique for Automated Root Cause Diagnosis of In-Production Failures]] ・ [[@2024__arXiv__Failure Diagnosis in Microservice Systems - A Comprehensive Survey and Analysis]] ### [OpsLens: Runbook-Free Root Cause Analysis via Multi-agent Collaboration and Tool Management](https://conf.researchr.org/track/ase-2026/ase-2026-industry-showcase) - 会場/年: ASE 2026 / 2026 - スコア: 88 - 収集元: 会議 - 査読: 査読済み - 推薦理由: 手順書に依存しない根本原因分析をマルチエージェント協調と道具管理で実現する提案であり、LLM による根本原因分析と SRE の自律化というコア関心の中心に当たる。手順書という人手資産への依存を外す方向は、実運用での適用範囲を広げる点で価値が高い。 - 関連: [[@2026__arXiv__Cloud-OpsBench - A Reproducible Benchmark for Agentic Root Cause Analysis in Cloud Systems]] ### [SCALER: LLM-based Cross-Modal Alignment for Microservice Root Cause Analysis](https://services.conferences.computer.org/2026/icws-program/) - 会場/年: ICWS 2026 / 2026 - スコア: 88 - 収集元: 会議 - 査読: 査読済み - 推薦理由: LLM を用いてモダリティ間の対応づけを行うマイクロサービス根本原因分析であり、コア関心の LLM による RCA に当たる。メトリクス・ログ・トレースの意味的整合という着想は多モーダル統合の好みに合致する。 - 関連: [[@2026__TOSEM__LLMRCA - Multilevel Root Cause Analysis for LLM Applications Using Multimodal Observability Data]] ・ [[@2023__ESEC-FSE__Nezha - Interpretable Fine-Grained Root Causes Analysis for Microservices on Multi-modal Observability Data]] ### [Smart Brain: Semantic Anomaly Detection for Operational Time Series in Large Scale Service Systems](https://conf.researchr.org/track/ase-2026/ase-2026-research-track) - 会場/年: ASE 2026 / 2026 - スコア: 88 - 収集元: 会議 - 査読: 査読済み - 推薦理由: 大規模サービスシステムの運用時系列に対する意味を伴う異常検知であり、コア関心の時系列異常検知と可観測性が交差する。意味情報で判定空間を絞る方針と、大規模実サービスを対象とする点の双方が評価の癖に合う。 - 関連: [[@2022__PVLDB__Anomaly Detection in Time Series - A Comprehensive Evaluation]] ・ [[@2024__SREcon24 EMEA__Anomaly Detection in Time Series from Scratch Using Statistical Analysis]] ### [CAVIAR: Disentangling Root Causes with an ICA-based VAE for Large-Scale Microservice Systems](https://kdd2026.kdd.org/papers/) - 会場/年: KDD 2026 / 2026 - スコア: 86 - 収集元: 会議 - 査読: 査読済み - 推薦理由: 独立成分分析に基づく変分自己符号化器で根本原因を分離する提案であり、大規模マイクロサービスの原因特定というコア関心の中心に当たる。混在する原因信号をもつれ解きの枠組みで分離する着想は、相関に頼る手法の限界を越える方向として評価できる。 - 関連: [[@2024__KDD__Microservice Root Cause Analysis with Limited Observability]] ・ [[@2020__IWQoS__Localizing Failure Root Causes in a Microservice through Causality Inference]] ### [ARMOR: A Robust Self-Supervised Framework for Root Cause Analysis in Microservices under Missing Modality](https://conf.researchr.org/track/ase-2026/ase-2026-research-track) - 会場/年: ASE 2026 / 2026 - スコア: 85 - 収集元: 会議 - 査読: 査読済み - 推薦理由: 一部のモダリティが欠落する状況でも頑健に働く自己教師ありの根本原因分析であり、コア関心に合致する。実運用ではテレメトリが常に揃わないという前提に立つ点が現場感覚に近い。 - 関連: [[@2023__ESEC-FSE__Nezha - Interpretable Fine-Grained Root Causes Analysis for Microservices on Multi-modal Observability Data]] ・ [[@2024__KDD__Microservice Root Cause Analysis with Limited Observability]] ### [AlarmClaw: Context-Enriched Alarm Management with Category-/Severity-Aware Incident Graphs](https://conf.researchr.org/track/ase-2026/ase-2026-industry-showcase) - 会場/年: ASE 2026 / 2026 - スコア: 85 - 収集元: 会議 - 査読: 査読済み - 推薦理由: 分類と重大度を考慮したインシデントグラフでアラートを文脈付けし管理する提案であり、アラート相関とインシデント管理というコア関心の中心にある。産業現場のアラート運用を対象とする点も加点材料である。 - 関連: [[@2024__CCGRID__AlertRCA - Causality Enhanced Graph Representation Learning for Alert-Based Root Cause Analysis]] ・ [[@2014__KDD__Unveiling Clusters of Events for Alert and Incident Management in Large-Scale Enterprise IT]] ### [DCRL: Dynamic Causal Representation Learning for Root Cause Localization in Microservice Systems](https://services.conferences.computer.org/2026/icws-program/) - 会場/年: ICWS 2026 / 2026 - スコア: 85 - 収集元: 会議 - 査読: 査読済み - 推薦理由: マイクロサービスの根本原因箇所特定を動的な因果表現学習で行う研究であり、コア関心に直接合致する。システムの構造が時間とともに変わる前提を扱う点は StaR と同じ問題意識であり、比較読みに向く。 - 関連: [[@2022__arXiv__CausalRCA - Causal Inference based Precise Fine-grained Root Cause Localization for Microservice Applications]] ・ [[@2024__DSN-S__Fault Localization Using Interventional Causal Learning for Cloud-Native Applications]] ### [HORATIO: Bridging Management and Analysis of Traces at Scale](https://www.semanticscholar.org/paper/a9588f5025be8f16bc1fb3c4e742999f260305fe) - 会場/年: SSDBM 2026: The International Conference on Scalable Scientific Data Management 2026 / 2026 - スコア: 85 - 収集元: メール - 査読: 査読済み - 推薦理由: テラバイト級の生トレースを変換せずその場で索引化・分析・物理クラスタ化する枠組みであり、テレメトリの保存と分析のスケーリングというコア関心に直結する。追加ストレージ約 1.01 倍で選択クエリを最大 75 倍高速化という、便益と代償を数値で示す提示の仕方が好みに合う。対象は HPC のプロファイリングトレースだが、分散トレーシングの保存基盤へ直接転用できる着想である。 - 関連: [[@2025__ISSTA__Tracezip - Efficient Distributed Tracing via Trace Compression]] ・ [[@2018__SoCC__Weighted Sampling of Execution Traces - Capturing More Needles and Less Hay]] ### [Metra: A Fault Localization Technique for Microservice Systems via the Fusion of Metrics and Traces](https://services.conferences.computer.org/2026/icws-program/) - 会場/年: ICWS 2026 / 2026 - スコア: 85 - 収集元: 会議 - 査読: 査読済み - 推薦理由: メトリクスとトレースを融合したマイクロサービスの障害箇所特定であり、単一信号でなく多層のテレメトリを統合する方針そのものである。コア関心の RCA と可観測性が重なる位置にある。 - 関連: [[@2022__IEEE CLOUD__Localizing and Explaining Faults in Microservices Using Distributed Tracing]] ・ [[@2025__TOSEM__Interpretable Failure Localization for Microservice Systems Based on Graph Autoencoder]] ### [TD-RCA: Topology-Aware and Dual-Perspective Decoupling for Root Cause Analysis in Microservices](https://iwqos2026.ieee-iwqos.org/detailed-technical-sessions) - 会場/年: IWQoS 2026 / 2026 - スコア: 85 - 収集元: 会議 - 査読: 査読済み - 推薦理由: トポロジを考慮し二視点で分離するマイクロサービス根本原因分析であり、コア関心に合致する。構造事前知識で探索空間を絞る方針と、視点を分離して評価する設計の双方が好みに合う。 - 関連: [[@2024__KDD__Microservice Root Cause Analysis with Limited Observability]] ・ [[@2026__arXiv__TopoEvo - A Topology-Aware Self-Evolving Multi-Agent Framework for Root Cause Analysis in Microservices]] ### [Towards Reliable Alarm Flood Reduction in Data Center via Grammar-Guided Verifiable LLM Reasoning](https://conf.researchr.org/track/ase-2026/ase-2026-industry-showcase) - 会場/年: ASE 2026 / 2026 - スコア: 85 - 収集元: 会議 - 査読: 査読済み - 推薦理由: データセンターのアラート洪水を、文法で制約した検証可能な LLM 推論により削減する提案である。アラート削減と疲労緩和というコア関心に直結し、出力を検証可能にする設計は推論プロセス志向の評価と相性が良い。 - 関連: [[@2026__OReilly__Observability Engineering 2E - Chapter 12 Acting On and Debugging SLO-Based Alerts]] ### [TuxBot: Semantic-Aware Online OS Tuning with LLMs](https://sigops.org/s/conferences/sosp/2026/accepted.html) - 会場/年: SOSP 2026 / 2026 - スコア: 85 - 収集元: 会議 - 査読: 査読済み - 推薦理由: LLM による OS のオンラインチューニングであり、柱 3 の学習ベース閉ループ制御に直接合致する。意味情報とドメイン知識で探索空間を絞る方針への共感とも一致し、SOSP 採択という水準も高い。 - 関連: [[@2026__arXiv__TuxBot - Semantic-Aware Online OS Tuning with Large Language Models]] ・ [[@2024__arXiv__OS Pre-trained Transformer - Predicting Query Latencies across Changing System Contexts]] ### [Waiting at the front door: Continuous monitoring of latency in the host network stack](https://conferences.sigcomm.org/imc/2026/accepted-papers/) - 会場/年: IMC 2026 / 2026 - スコア: 85 - 収集元: 会議 - 査読: 査読済み - 推薦理由: ホストのネットワークスタック内部の遅延を継続的に監視する研究であり、低オーバーヘッドな常時計測というコア関心に正面から合致する。アプリとネットワークの境界にある遅延をホスト側で分解する視点は、多層テレメトリ統合の好みにも合う。IMC の計測研究として実測データに基づく評価が期待できる。 - 関連: [[@2019__SREcon19 Americas__Latency SLOs Done Right]] ・ [[@2024__arXiv__OS Pre-trained Transformer - Predicting Query Latencies across Changing System Contexts]] ### [When Does AI Actually Help in Incident Response? Identifying Good First Messages in Cloud Service Incidents](https://conf.researchr.org/track/ase-2026/ase-2026-research-track) - 会場/年: ASE 2026 / 2026 - スコア: 85 - 収集元: 会議 - 査読: 査読済み - 推薦理由: クラウドインシデントにおいて AI の介入が実際に効く場面を特定する実証研究であり、インシデント管理の自律化というコア関心に当たる。効果があると主張するのでなく効く条件を切り分ける姿勢は、成果志向でなくプロセス志向という評価の癖に合致する。 - 関連: [[@2025__SRE NEXT 2025__Rethinking Incident Response - Context-Aware AI in Practice]] ・ [[@2020__ASE__How Incidental are the Incidents - Characterizing and Prioritizing Incidents for Large-Scale Online Service Systems]] ### [DualLane: Fast and Reliable LLM Agents for Interactive AIOps via Dual-Path Planning](https://kdd2026.kdd.org/papers/) - 会場/年: KDD 2026 / 2026 - スコア: 84 - 収集元: 会議 - 査読: 査読済み - 推薦理由: 対話型 AIOps 向けに高速経路と慎重経路を使い分ける二重計画で LLM エージェントの応答性と信頼性を両立させる研究であり、SRE エージェントというコア関心の直撃である。速さと確実性のトレードオフを構造として明示する設計は、コストと便益を数値で示す好みにも合う。 - 関連: [[@2026__FSE Companion__LLM Agents for AIOps in Kubernetes - An Industrial Experience Report with Red Hat OpenShift]] ・ [[@2026__arXiv__AgentTether - Graph-Guided Diagnosis and Runtime Intervention for Reliable LLM Agent Operations]] ### [LMID: A Comprehensive Multimodal Dataset for Failure Prediction in Cloud Computing Systems](https://kdd2026.kdd.org/papers/) - 会場/年: KDD 2026 / 2026 - スコア: 84 - 収集元: 会議 - 査読: 査読済み - 推薦理由: クラウドシステムの障害予測に向けた包括的なマルチモーダルデータセットであり、予兆検知と AIOps の評価設計という二つの関心に同時に当たる。単一信号でなくメトリクス・ログ・トレースなど複数モダリティを統合する方向は方法論の好みに合致する。データセット公開型の研究は再現性の面でも価値が高い。 - 関連: [[@2025__arXiv__LogDB - Multivariate Log-based Failure Diagnosis for Distributed Databases]] ・ [[@2019__WWW__Outage Prediction and Diagnosis for Cloud Service Systems]] ### [CAPMix: Robust KPI Anomaly Detection for AIOps in Noisy and Dynamic Environments](https://conf.researchr.org/track/ase-2026/ase-2026-industry-showcase) - 会場/年: ASE 2026 / 2026 - スコア: 82 - 収集元: 会議 - 査読: 査読済み - 推薦理由: 雑音が多く変動する実環境における KPI 異常検知の頑健化であり、運用テレメトリの異常検知というコア関心に直接該当する。理想化されたベンチマークではなく実運用の雑音を前提としている点が加点材料になる。 - 関連: [[@2022__ICWS__TS-InvarNet - Anomaly Detection and Localization based on Tempo-spatial KPI Invariants in Distributed Services]] ・ [[@2025__arXiv__KPIRoot+ - An Efficient Integrated Framework for Anomaly Detection and Root Cause Analysis in Large-Scale Cloud Systems]] ### [Few-shot Multimodal Anomaly Detection via Dynamic Intra-modal Sparsity Attention and Quality-aware Cross-modal Fusion in Microservice System](https://kdd2026.kdd.org/papers/) - 会場/年: KDD 2026 / 2026 - スコア: 82 - 収集元: 会議 - 査読: 査読済み - 推薦理由: マイクロサービスにおけるモダリティ内スパース注意とモダリティ間の品質考慮融合による少数事例異常検知であり、多モーダル統合の好みに正面から合致する。品質に応じて融合重みを変える設計は、信号ごとに信頼度が異なる実運用テレメトリの性質を捉えている。 - 関連: [[@2024__ASE__Giving Every Modality a Voice in Microservice Failure Diagnosis via Multimodal Adaptive Optimization]] ・ [[@2021__ESEC-FSE__Identifying Bad Software Changes via Multimodal Anomaly Detection]] ### [LogNexus: Effective Log Compression via Unified Redundancy Encoding](https://conf.researchr.org/track/issta-2026/issta-2026-research-papers) - 会場/年: ISSTA 2026 / 2026 - スコア: 82 - 収集元: 会議 - 査読: 査読済み - 推薦理由: 冗長性を統一的に符号化してログを圧縮する。テレメトリ量の削減という中核課題に直接効き、圧縮率と処理コストのトレードオフが数値で示される見込みが高い。Tracezip の系列として継続して追う価値がある。 - 関連: [[@2023__ICSE__LogReducer - Identify and Reduce Log Hotspots in Kernel on the Fly]] ・ [[@2025__ISSTA__Tracezip - Efficient Distributed Tracing via Trace Compression]] ### [Ranking Anomalous Subsequences in Large Scale Streaming Data for Efficient Triage](https://kdd2026.kdd.org/papers/) - 会場/年: KDD 2026 / 2026 - スコア: 82 - 収集元: 会議 - 査読: 査読済み - 推薦理由: 大規模ストリーミングデータ上で異常部分列を順位付けし、トリアージの効率化につなげる提案である。時系列異常検知とアラート優先順位付けというコア関心の交点にあり、大規模運用データを前提とする点も加点材料である。 - 関連: [[@2017__KDD__Anomaly Detection in Streams with Extreme Value Theory]] ・ [[@2025__VLDB__RCRank - Multimodal Ranking of Root Causes of Slow Queries in Cloud Database Systems]] ### [Tracing the Invisible: Semantic Message Flow Discovery via Data Contracts in Real-World Distributed Systems](https://conf.researchr.org/track/ase-2026/ase-2026-research-track) - 会場/年: ASE 2026 / 2026 - スコア: 82 - 収集元: 会議 - 査読: 査読済み - 推薦理由: データ契約を手がかりに実運用の分散システムで意味的なメッセージフローを発見する研究であり、非同期経路の可視化という分散トレーシングの難所を扱う。構造的なドメイン知識で探索空間を絞る発想と、実システムでの検証という点の双方が好みに合致する。 - 関連: [[@2022__IEEE CLOUD__Localizing and Explaining Faults in Microservices Using Distributed Tracing]] ### [Concordia: JIT-Compiled Persistent-Kernel Checkpointing for Fault-Tolerant LLM Inference](https://www.semanticscholar.org/paper/f6ed1160493df8dc3ba65232c342afcfd06a1f82) - 会場/年: arXiv (Cornell University) / 2026 - スコア: 80 - 収集元: メール - 査読: arXiv のみ - 推薦理由: GPU 常駐の永続カーネルを基盤に据えた LLM 推論の耐障害チェックポイントであり、学習・推論基盤の高速リカバリというコア関心に該当する。PTX/SASS 水準でフレームワーク境界の下に計装を差し込む設計は、低オーバーヘッド動的計装の関心とも重なる。 - 関連: [[@2026__OSDI__SDCs in the Wild - Characterizing and Diagnosing SDC-defective GPUs in Production LLM Training]] ・ [[@2025__arXiv__LMCache - An Efficient KV Cache Layer for Enterprise-Scale LLM Inference]] ### [Discovering Performance Archetypes: Critical-Path-Aware Pattern Analysis and Regression Detection](https://conf.researchr.org/track/ase-2026/ase-2026-research-track) - 会場/年: ASE 2026 / 2026 - スコア: 80 - 収集元: 会議 - 査読: 査読済み - 推薦理由: クリティカルパスを考慮した性能パターンの類型化と回帰検知であり、分散システムの性能診断というコア関心に合致する。CRISP 系のクリティカルパス解析を発展させる方向で、トレース由来の構造情報を活用する点も好ましい。 - 関連: [[@2022__USENIX ATC__CRISP - Critical Path Analysis of Large-Scale Microservice Architectures]] ・ [[@2015__CSUR__Performance Anomaly Detection and Bottleneck Identification]] ### [Falconf: Configuration Error Diagnosis via Log Sequence Learning and Automated Misconfiguration Injections](https://icdcs2026.icdcs.org/program/main-technical-sessions/) - 会場/年: ICDCS 2026 / 2026 - スコア: 80 - 収集元: 会議 - 査読: 査読済み - 推薦理由: ログ系列の学習と設定誤りの自動注入を組み合わせて設定起因障害を診断する。ログベース診断と障害注入という関心が重なり、注入によるラベル生成は評価設計の観点でも参考になる。 - 関連: [[@2024__DSN-S__Fault Localization Using Interventional Causal Learning for Cloud-Native Applications]] ・ [[@2011__SOSP__An Empirical Study on Configuration Errors in Commercial and Open Source Systems]] ### [In a Streaming World, Should You Stand Still? A Comprehensive Benchmark of Anomaly Detection in Streams](https://kdd2026.kdd.org/papers/) - 会場/年: KDD 2026 / 2026 - スコア: 80 - 収集元: 会議 - 査読: 査読済み - 推薦理由: ストリーミング環境における異常検知手法の包括的ベンチマークであり、時系列異常検知と評価設計という二つの関心が交差する。オンライン更新の要否を問う設計は、運用テレメトリでの継続学習コストの判断材料になる。 - 関連: [[@2017__KDD__Anomaly Detection in Streams with Extreme Value Theory]] ・ [[@2022__PVLDB__Anomaly Detection in Time Series - A Comprehensive Evaluation]] ### [Mantis: Decoding HPC Telemetry Data for Robust System Prediction](https://dipsa-qub.github.io/ICS2026-webpage/program/program.html) - 会場/年: ICS 2026 / 2026 - スコア: 80 - 収集元: 会議 - 査読: 査読済み - 推薦理由: 大規模 HPC システムのテレメトリを解読して将来のシステム状態を予測する。テレメトリ工学と予兆検知というコア関心に直に当たり、実運用規模のデータを扱う見込みが高い。 - 関連: [[@2025__SC__Fine-grained Automated Failure Management for Extreme-Scale GPU Accelerated Systems]] ・ [[@2025__APNET__Forewarned is Forearmed - Joint Prediction and Classification of Optical Transceiver Failures in Large-Scale LLM Training Clusters]] ### [PHOENIX: Resilient LLM Training with Hot-Swapping via Zero-Overhead Checkpoint](https://www.semanticscholar.org/paper/bd792779baf8a4167f30c6b82c854b99dc8daa39) - 会場/年: arXiv (Cornell University) / 2026 - スコア: 80 - 収集元: メール - 査読: arXiv のみ - 推薦理由: 故障ノードをジョブを止めずに交換するホットスワップで LLM 学習を復旧させ、無障害時のオーバーヘッドを実質ゼロに抑える。無障害時コストと復旧遅延のパレート境界を明示的に扱う点が、便益と代償を数値で示す研究への好みに合致する。 - 関連: [[@2026__arXiv__ReCoVer - Resilient LLM Pre-Training System via Fault-Tolerant Collective and Versatile Workload]] ・ [[@2026__SIGCOMM__Connecting 100K+ GPUs - Building the Communication Stack for Large-Scale LLM Training]] ### [SG-LR QoS: Sparse-Gating and Layered Low-Rank Based Lightweight QoS Monitoring for Multi-tenant Clouds](https://services.conferences.computer.org/2026/icws-program/) - 会場/年: ICWS 2026 / 2026 - スコア: 80 - 収集元: 会議 - 査読: 査読済み - 推薦理由: 多テナントクラウドの品質監視を疎なゲーティングと階層的低ランク近似で軽量化する研究であり、テレメトリ量の削減というコア関心に直撃する。監視精度と計算コストのトレードオフを明示する構図は方法論の好みにも合致する。 - 関連: [[@2024__ASE__SLIM - A scalable and interpretable light-weight fault localization algorithm for imbalanced data in microservice]] ・ [[@2026__TOSEM__LLMRCA - Multilevel Root Cause Analysis for LLM Applications Using Multimodal Observability Data]] ### [StepFinder: A Temporal Semantic Framework for Failure Attribution in Multi-Agent Systems](https://kdd2026.kdd.org/papers/) - 会場/年: KDD 2026 / 2026 - スコア: 80 - 収集元: 会議 - 査読: 査読済み - 推薦理由: マルチエージェント系の障害をどの手順のどの時点に帰属させるかを時間的かつ意味的に特定する枠組みであり、箇所特定というコア関心に直結する。最終結果でなく過程を分解して評価する構図も好みに合う。 - 関連: [[@2025__arXiv__StepFly - Agentic Troubleshooting Guide Automation for Incident Diagnosis]] ・ [[@2024__ASE__The Potential of One-Shot Failure Root Cause Analysis - Collaboration of the Large Language Model and Small Classifier]] ### [TORCL: Lightweight Root Cause Localization with Temporal-Order-Aware Ranking in Microservices](https://services.conferences.computer.org/2026/icws-program/) - 会場/年: ICWS 2026 / 2026 - スコア: 80 - 収集元: 会議 - 査読: 査読済み - 推薦理由: 時間順序を考慮した順位付けによるマイクロサービスの根本原因箇所特定であり、コア関心に直接該当する。軽量性を前面に出しており、診断機構自体のオーバーヘッドを抑える設計は運用適用の観点で好ましい。 - 関連: [[@2026__Elsevier__MicroIRC - Instance-level Root Cause Localization for Microservice Systems]] ・ [[@2020__IWQoS__Localizing Failure Root Causes in a Microservice through Causality Inference]] ### [Taming Multi-Dimensional Tail: A Factorized Tail-Interaction Framework for Hyperscale Cloud Workload Forecasting](https://kdd2026.kdd.org/papers/) - 会場/年: KDD 2026 / 2026 - スコア: 80 - 収集元: 会議 - 査読: 査読済み - 推薦理由: ハイパースケールのクラウドワークロード予測において多次元の裾を因子分解して扱う研究であり、実運用規模での評価という好みに強く合致する。裾の挙動は容量計画と自動スケーリングの成否を左右するため、柱 3 に直接効く。 - 関連: [[@2026__SIGCOMM__Rethinking Cloud Optimization - Volatility-Driven for Better Outcomes]] ### [TuneAgent: Agentic Operating System Kernel Tuning with Reinforcement Learning](https://kdd2026.kdd.org/papers/) - 会場/年: KDD 2026 / 2026 - スコア: 80 - 収集元: 会議 - 査読: 査読済み - 推薦理由: 強化学習を用いたエージェント型のオペレーティングシステム核チューニングであり、学習ベースの自動チューニングというコア関心に合致する。低レイヤの計測結果を制御に還元する構図も関心に近い。 - 関連: [[@2025__arXiv__The Landscape of Agentic Reinforcement Learning]] ## 注目 スコア 60–79。138 件。件数が多いため一覧表形式とし、推薦理由は 1 行に要約した。 | スコア | 会場 | タイトル | 要点 | 収集元 | | --- | --------------------------------------------------------------------------------------------- | ---------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- | -------------------------------------------------- | ------ | | 79 | ICS 2026 | [DA-MLAD: Drift-Decomposed Meta-Learning for Continual Log Anomaly Detection in Supercomputing Systems](https://dipsa-qub.github.io/ICS2026-webpage/program/program.html) | 分布ずれを分解する継続的ログ異常検知。スーパーコンピュータの実ログが対象で概念ドリフトを正面から扱う | 会議 | | 79 | ASE 2026 | [Disassembly-Augmented Root Cause Analysis for C/C++ Segmentation Faults in SAP HANA](https://conf.researchr.org/track/ase-2026/ase-2026-industry-showcase) | 逆アセンブル情報を用いた C/C++ セグメンテーション違反の RCA。SAP HANA の産業事例 | 会議 | | 78 | SOSP 2026 | [A Few GPUs, A Whole Lotta Scale: Faithful LLM Training Emulation with CrystalLLM](https://sigops.org/s/conferences/sosp/2026/accepted.html) | 少数 GPU で大規模 LLM 学習の挙動を忠実に再現するエミュレーション。安価な反復実験の基盤 | 会議 | | 78 | KDD 2026 | [CAAD: Causality-Aware Multivariate Time Series Anomaly Detection via Multi-Scale Alignment and Structural Causal Consistency](https://kdd2026.kdd.org/papers/) | 構造的因果整合性を制約に用いた多変量時系列異常検知。異常検知と因果の交点 | 会議 | | 78 | ASE 2026 | [Expanding Guard Reach: Online Learning-Driven Incremental Kernel Panic Diagnosis](https://conf.researchr.org/track/ase-2026/ase-2026-industry-showcase) | オンライン学習でカーネルパニック診断の適用範囲を漸進的に広げる。実運用クラッシュログが対象 | 会議 | | 78 | ISCA 2026 | [Lit Silicon: A Case Where Thermal Imbalance Couples Concurrent Execution in Multiple GPUs](https://iscaconf.org/isca2026/program/) | 複数 GPU の熱的不均衡が並行実行を結合させる性能異常の機序解明。物理層から上位への伝搬経路を示す | 会議 | | 78 | SOSP 2026 | [RPCShield: Defending Microservices Against Cascading Failures](https://sigops.org/s/conferences/sosp/2026/accepted.html) | マイクロサービスのカスケード障害を RPC 層で防御。障害伝搬と緩和を扱う | 会議 | | 78 | IWQoS 2026 | [SE-STGAE: A Shapelet-Enhanced Spatiotemporal Graph Autoencoder for Early Anomaly Detection](https://iwqos2026.ieee-iwqos.org/detailed-technical-sessions) | シェイプレット強化の時空間グラフ自己符号化器による早期異常検知。予兆検知に該当 | 会議 | | 78 | IWQoS 2026 | [SketchPlan: Full-Visibility Sketch-Based Telemetry with Limited Programmable Switch Coverage](https://iwqos2026.ieee-iwqos.org/detailed-technical-sessions) | プログラマブルスイッチの部分配備でも全体可視性を確保するスケッチ型テレメトリ | 会議 | | 78 | ICS 2026 | [TenProf: A Tensor-Centric Profiler for Deep Learning Workload Analysis and Optimization](https://dipsa-qub.github.io/ICS2026-webpage/program/program.html) | テンソル単位で捉える深層学習プロファイラ。カーネル単位でない観測の切り口 | 会議 | | 78 | KDD 2026 | [TimeRadar: A Domain-Rotatable Foundation Model for Time Series Anomaly Detection](https://kdd2026.kdd.org/papers/) | ドメイン回転で適応する時系列異常検知の基盤モデル。ゼロショット適用で検知器作り込みを減らす | 会議 | | 78 | ICS 2026 | [Wattchmen: Watching the Wattchers – High Fidelity, Flexible GPU Energy Modeling](https://dipsa-qub.github.io/ICS2026-webpage/program/program.html) | GPU 電力を高忠実度かつ柔軟にモデル化。電力という代償の定量化 | 会議 | | 77 | ISCA 2026 | [Enabling Continuous, In-Field Introspection: The Programmable IPU Architecture](https://iscaconf.org/isca2026/program/) | 実地で継続的な内部観測を可能にするプログラマブル IPU。計装コストをホスト CPU から切り離す | 会議 | | 77 | ISCA 2026 | [Observability-aided GPU Memory Oversubscription](https://iscaconf.org/isca2026/program/) | 観測情報を入力に GPU メモリの超過割り当てを制御。観測を最適化の一次入力に据える | 会議 | | 77 | KDD 2026 | [ReCATS: Replay-Free Continual Anomaly Detection for Non-Stationary Multivariate Time Series](https://kdd2026.kdd.org/papers/) | 再生バッファ不要の継続的異常検知。過去データを保持せず保存コストを抑える | 会議 | | 76 | ICDCS 2026 | [FloodGuard: A Prediction-Control Closed Loop for Mitigating Cold-Start Floods in Cloud Services](https://icdcs2026.icdcs.org/program/main-technical-sessions/) | コールドスタート殺到を予測と制御の閉ループで緩和。予兆検知から緩和まで一続き | 会議 | | 76 | SOSP 2026 | [Observability-Driven Tiered Memory for Colocated Workloads](https://sigops.org/s/conferences/sosp/2026/accepted.html) | 観測データを駆動源に同居ワークロードの階層メモリを制御。計測から制御への閉ループ | 会議 | | 75 | IWQoS 2026 | [Breaking the Accuracy-Latency Trade-Off in Sketch Compression for Network Measurement](https://iwqos2026.ieee-iwqos.org/detailed-technical-sessions) | ネットワーク計測におけるスケッチ圧縮の精度と遅延の二律背反を崩す | 会議 | | 75 | KDD 2026 | [State Machine Guided Multi-Relational Synthetic Data from Logs for Anomaly Detection](https://kdd2026.kdd.org/papers/) | ログからの状態機械誘導による合成データ生成。ログ異常検知の学習データ不足に効く | 会議 | | 75 | SOSP 2026 | [Taming Inference Workloads at Global Scale: Foundation Model Serving in SystemX](https://sigops.org/s/conferences/sosp/2026/accepted.html) | グローバル規模の基盤モデルサービングの実運用報告。AI インフラ運用の実態資料 | 会議 | | 75 | ICWS 2026 | [Transferable Unsupervised Time Series Anomaly Detection with Mixture-of-Experts for Web Systems](https://services.conferences.computer.org/2026/icws-program/) | 専門家混合による転移可能な教師なし時系列異常検知。サービスごとの学習コストを抑える | 会議 | | 74 | KDD 2026 | [A Physics-Aware Framework for Short-Term GPU Power Forecasting of AI Data Centers](https://kdd2026.kdd.org/papers/) | 物理法則を組み込んだ AI データセンターの短期 GPU 電力予測。構造事前知識の活用 | 会議 | | 74 | KDD 2026 | [Beyond Holistic Models: Systematic Component-level Benchmarking of Deep Multivariate Time-Series Forecasting](https://kdd2026.kdd.org/papers/) | 深層多変量時系列予測を構成要素ごとに分解評価するベンチマーク。推論プロセス志向の評価設計 | 会議 | | 74 | ICS 2026 | [Not All Errors Are Equal: A Systematic Study of Error Propagation in Large Language Model Inference](https://dipsa-qub.github.io/ICS2026-webpage/program/program.html) | LLM 推論におけるエラー伝搬の系統的調査。誤りが等価でないという主張は影響伝搬に接続 | 会議 | | 74 | ASE 2026 | [Panorama: Unveiling Latent Dependencies for Microservice Autoscaling via Meta-learning](https://conf.researchr.org/track/ase-2026/ase-2026-research-track) | メタ学習でマイクロサービス間の潜在依存を掘り起こし自動スケーリングに活かす | 会議 | | 74 | ASE 2026 | [ReLog: Execution-Aware Logging with Runtime Feedback for LLM-Oriented Debugging](https://conf.researchr.org/track/ase-2026/ase-2026-research-track) | 実行時の反応を取り込んでログ出力そのものを設計。何をどれだけ記録するかの上流を扱う | 会議 | | 72 | ICWS 2026 | [A Retrieval-Augmented Generation-driven Auto-scaling Method for Microservices in Cloud-native Environments](https://services.conferences.computer.org/2026/icws-program/) | 検索拡張生成によるマイクロサービス自動スケーリング。過去の運用知識を制御に反映 | 会議 | | 72 | ICS 2026 | [Agile QoS-aware Dynamic Power Management with eBPF Governors](https://dipsa-qub.github.io/ICS2026-webpage/program/program.html) | eBPF による省電力ガバナ。eBPF を観測でなくカーネル内閉ループ制御に用いる | 会議 | | 72 | IMC 2026 | [Behavioral Consistency and Transparency Analysis on Large Language Model API Gateways](https://conferences.sigcomm.org/imc/2026/accepted-papers/) | LLM API ゲートウェイの挙動一貫性と透明性の第三者計測。自己申告に依らない実測 | 会議 | | 72 | KDD 2026 | [Bridging Classification and Reconstruction: Cooperative Time Series Anomaly Detection](https://kdd2026.kdd.org/papers/) | 分類ベースと再構成ベースの時系列異常検知を協調させる枠組み | 会議 | | 72 | ASE 2026 | [Bridging the Gap between Intent and Impact: An Empirical Study of GPU Optimizations in Deep Learning Frameworks](https://conf.researchr.org/track/ase-2026/ase-2026-research-track) | 深層学習フレームワークの GPU 最適化が意図どおり効いているかの実証。意図と実測の乖離を測る | 会議 | | 72 | ICWS 2026 | [CogGuard: Cognitive and Operational Profiling for Proactive Warning in Edge Intelligent Services](https://services.conferences.computer.org/2026/icws-program/) | エッジ知的サービスの予兆警告をプロファイリングで実現。プロアクティブな障害管理 | 会議 | | 72 | World Wide Web | [DAIMS: a distributed multi-agent framework for explainable anomaly detection in web services using machine learning and semantic reasoning](https://www.semanticscholar.org/paper/edfe17feaad2e4b5ea7f07b6316ede9f6ba1036d) | 分散マルチエージェントと意味推論による説明可能な異常検知 | メール | | 72 | ASE 2026 | [Deployment Risk Assessment using Diff-Aware Features: A Case Study at Prime Video](https://conf.researchr.org/track/ase-2026/ase-2026-industry-showcase) | Prime Video の実運用で変更差分からデプロイのリスクを見積もる事例 | 会議 | | 72 | KDD 2026 | [LatentFlow: Discovering Latent Continuous Dynamics across Channels for Multivariate Time Series Anomaly Detection](https://kdd2026.kdd.org/papers/) | チャネル間の潜在連続ダイナミクスを捉える多変量時系列異常検知 | 会議 | | 72 | ISCA 2026 | [PIPEWEAVE : Synergizing Analytical and Learning Models for Unified GPU Performance Prediction](https://iscaconf.org/isca2026/program/) | 解析モデルと学習モデルを統合した GPU 性能予測。AI インフラの性能可視化 | 会議 | | 72 | HPDC '26: 35th International Symposium on High-Performance Parallel and Distributed Computing | [Robust Ensemble Log Anomaly Detection for High-performance Parallel and Distributed Computing Systems](https://www.semanticscholar.org/paper/4daa71cc6ef442270946c997d0cae6aa098a1754) | 大規模並列分散基盤向けのストリーミング型ログ異常検知アンサンブル。閾値の自動較正まで踏み込む | メール | | 72 | KDD 2026 | [SEA-FGT: Frequency-Guided Transformer with Semantic Expert Augment for Time Series Anomaly Detection](https://kdd2026.kdd.org/papers/) | 周波数領域と意味的専門家機構を組み合わせた時系列異常検知 | 会議 | | 72 | ICWS 2026 | [ST-RCA: A Task-Oriented Causal Framework for Robust Spatio-Temporal Root Cause Analysis in Complex IoT Systems](https://services.conferences.computer.org/2026/icws-program/) | 時空間の因果枠組みによる RCA。対象は IoT だが手法の枠組みは転用可能 | 会議 | | 72 | ASE 2026 | [SimServing: Native-Execution Simulation for Evolution-Resilient LLM Serving Configuration Tuning](https://conf.researchr.org/track/ase-2026/ase-2026-research-track) | LLM サービングの設定チューニングを実行相当のシミュレーションで支える | 会議 | | 72 | IWQoS 2026 | [Taming GPU Inference Variability with Risk-Aware Scheduling and Dynamic Kernel Segments](https://iwqos2026.ieee-iwqos.org/detailed-technical-sessions) | GPU 推論の性能ばらつきをリスク認識スケジューリングで抑える。変動の定量化部分は観測に転用可 | 会議 | | 72 | ASE 2026 | [What Breaks When LLMs Code? Characterizing Operational Safety Failures of Agentic Code Assistants](https://conf.researchr.org/track/ase-2026/ase-2026-research-track) | エージェント型コード補助が引き起こす運用安全性障害の類型化。失敗モード辞書として有用 | 会議 | | 70 | KDD 2026 | [DAMR: Dual Adaptive Multi-Head Representation Learning for Multivariate Time Series Anomaly Detection](https://kdd2026.kdd.org/papers/) | 適応的な多頭表現学習による多変量時系列異常検知 | 会議 | | 70 | IWQoS 2026 | [DNSMedic: Automated Precision Diagnosis and Repair Guidance for DNS Misconfigurations](https://iwqos2026.ieee-iwqos.org/detailed-technical-sessions) | DNS 設定ミスの精密診断と修復指針の自動生成。原因特定から復旧提案まで | 会議 | | 70 | KDD 2026 | [G-STAR: Graph-based Scheduling with Trace-driven Adaptive Routing for Industrial LLM-based Multi-Agent Systems](https://kdd2026.kdd.org/papers/) | 産業規模の LLM マルチエージェント基盤をトレース駆動でスケジューリング | 会議 | | 70 | ASE 2026 | [InfraAPR: An Agent-Based Repair Framework for Cloud Infrastructure-as-Code Programs](https://conf.researchr.org/track/ase-2026/ase-2026-industry-showcase) | エージェントによるクラウド構成コードの自動修復。セルフヒーリングの系譜 | 会議 | | 70 | SOSP 2026 | [It's the Kernel's Fault! Custom Page Fault Handling With bpf_fault](https://sigops.org/s/conferences/sosp/2026/accepted.html) | ページフォールト処理を BPF で拡張可能にする。カーネル層の観測と制御の境界を広げる | 会議 | | 70 | KDD 2026 | [LEFT: Learnable Fusion of Tri-view Tokens for Unsupervised Time Series Anomaly Detection](https://kdd2026.kdd.org/papers/) | 三視点トークンの学習可能な融合による教師なし時系列異常検知 | 会議 | | 70 | KDD 2026 | [MEMTS: Internalizing Domain Knowledge via Parameterized Memory for Retrieval-Free Domain Adaptation of Time Series Foundation Models](https://kdd2026.kdd.org/papers/) | 時系列基盤モデルを検索に頼らず母数化された記憶でドメイン適応させる | 会議 | | 70 | ISCA 2026 | [MTIA 300: Meta's First Training Chip Featuring Built-in NICs and Collective Offloading Engines](https://www.semanticscholar.org/paper/f497d93ab2c8ceba0e59f09ad4886459a3b17d90) | Meta 初の学習用チップ。集合通信の卸下機構を内蔵。性能値は自社申告のため割り引いて読む | メール+会議 | | 70 | SOSP 2026 | [Multi-LLM Serving at Production Scale](https://sigops.org/s/conferences/sosp/2026/accepted.html) | 実運用規模での複数 LLM 同時提供。資源競合と性能異常の知見が期待できる | 会議 | | 70 | ICWS 2026 | [Resource-Efficient Online Reliability Anomaly Detection for Edge Services via Deep Reinforcement Learning](https://services.conferences.computer.org/2026/icws-program/) | エッジサービスの信頼性異常をオンライン検知しつつ資源消費を抑える | 会議 | | 70 | IWQoS 2026 | [Square: Towards Sharing Internet-Scale GPU Clouds Fairly and Efficiently](https://iwqos2026.ieee-iwqos.org/detailed-technical-sessions) | インターネット規模の GPU クラウドを公平かつ効率的に共有する仕組み | 会議 | | 70 | 2026 IEEE 26th International Symposium on Cluster, Cloud and Internet Computing (CCGrid) | [TelePod: Live Migration for Stateful Containers](https://www.semanticscholar.org/paper/d6a4ba533eb721492e7f5747f3f35c2a7d76b03c) | 状態保持したままのコンテナ生移送。無停止退避は自己修復の基盤技術 | メール | | 70 | ICWS 2026 | [Toward Effective Resource Prediction for Microservices via Heterogeneity Exploration Fully Investigating Services, Containers, and Nodes](https://services.conferences.computer.org/2026/icws-program/) | サービス・コンテナ・ノードの三層を横断したマイクロサービス資源予測 | 会議 | | 70 | ISCA 2026 | [Understanding Inference Scaling for LLMs: Bottlenecks, Trade-offs, and Performance Principles](https://iscaconf.org/isca2026/program/) | LLM 推論スケーリングをボトルネックとトレードオフの観点から原理化 | 会議 | | 68 | SOSP 2026 | [Batched in Back: Characterizing and Optimizing Offline LLM Inference in Production with ACDC](https://sigops.org/s/conferences/sosp/2026/accepted.html) | 実運用における一括処理型 LLM 推論の特性解析。本番実測に基づく特性把握 | 会議 | | 68 | IWQoS 2026 | [Beagle: Auto-tuning Performance Diagnosis for Live Video Streaming](https://iwqos2026.ieee-iwqos.org/detailed-technical-sessions) | ライブ動画配信の性能診断を自動チューニングと結び付ける | 会議 | | 68 | ICDCS 2026 | [Coordinating Traffic at Group Granularity for Tail Latency Control in Service Meshes](https://icdcs2026.icdcs.org/program/main-technical-sessions/) | サービスメッシュでトラフィックをグループ粒度に調整しテール遅延を抑える | 会議 | | 68 | IWQoS 2026 | [GaLB: Gap-Aware Load Balancing for High-Performance AI Training in RDMA Networks](https://iwqos2026.ieee-iwqos.org/detailed-technical-sessions) | RDMA ネットワークにおける AI 学習向けの間隙認識型負荷分散 | 会議 | | 68 | ICS 2026 | [Harnessing MPI mutations for AI error detection](https://dipsa-qub.github.io/ICS2026-webpage/program/program.html) | MPI の変異注入により AI 学習中の誤りを検知。TrainCheck と同じ問題意識 | 会議 | | 68 | ASE 2026 | [HyTri: Hybrid Triage of CI Failures via LLM Guided Semantic Reasoning and Change Attribution](https://conf.researchr.org/track/ase-2026/ase-2026-research-track) | 継続的統合の失敗を意味的推論と変更帰属の併用で分類 | 会議 | | 68 | KDD 2026 | [LF-Filter: Differentiated Processing for Intra-Flow Packet Delay Monitoring in High-Speed Data Streams](https://kdd2026.kdd.org/papers/) | 高速データストリームでフロー内パケット遅延を差別化処理して監視 | 会議 | | 68 | IWQoS 2026 | [Learning from Experience: Real-time and Efficient Network Measurement on Programmable Switches](https://iwqos2026.ieee-iwqos.org/detailed-technical-sessions) | プログラマブルスイッチ上での実時間かつ資源効率のよいネットワーク計測 | 会議 | | 68 | ISSTA 2026 | [Profiling-Guided Bayesian Optimization of JVM Configurations](https://conf.researchr.org/track/issta-2026/issta-2026-research-papers) | プロファイリング結果を手がかりにベイズ最適化で JVM 設定を調整 | 会議 | | 68 | IWQoS 2026 | [RiCE: Precise Remote In-Network Congestion Elimination In Inter-Datacenter RDMA Networks](https://iwqos2026.ieee-iwqos.org/detailed-technical-sessions) | データセンタ間 RDMA の遠隔輻輳をインネットワークで精密に除去 | 会議 | | 68 | ISCA 2026 | [Scalable Synthesis of Distributed LLM Workloads Through Symbolic Tensor Graphs](https://iscaconf.org/isca2026/program/) | 記号的テンソルグラフから分散 LLM ワークロードを合成。実機なしの性能解析 | 会議 | | 68 | ICWS 2026 | [TFLLM4Load: Time-Frequency Enhanced LLM-enabled Framework for Multi-Scale Workload Prediction with External Forecasting Priors](https://services.conferences.computer.org/2026/icws-program/) | 時間周波数情報と外部予測事前情報を組み合わせた LLM ワークロード予測 | 会議 | | 68 | KDD 2026 | [TemporalBench: A Benchmark for Evaluating LLM-Based Agents on Contextual and Event-Informed Time Series Tasks](https://kdd2026.kdd.org/papers/) | 文脈と事象を織り込んだ時系列課題における LLM エージェントの評価基準 | 会議 | | 68 | ASE 2026 | [Understanding Bugs in Modern Agentic Frameworks: A Study of Symptoms, Root Causes, and Triggering Conditions](https://conf.researchr.org/track/ase-2026/ase-2026-research-track) | エージェントフレームワークの不具合を症状・根本原因・発現条件の三層で整理 | 会議 | | 68 | IMC 2026 | [When BGP Goes MAD: Multi-Scale Anomaly Detection in Internet Routing](https://conferences.sigcomm.org/imc/2026/accepted-papers/) | インターネット経路制御における多尺度の異常検知。IMC の実測評価 | 会議 | | 66 | SOSP 2026 | [Beyond Utilization: Energy-Conscious GPU Sharing for Inference Serving](https://sigops.org/s/conferences/sosp/2026/accepted.html) | 利用率という指標を超えエネルギーを軸に GPU 共有を設計。指標そのものを問い直す | 会議 | | 66 | SOSP 2026 | [DDB: Source-Level Interactive Debugging for Distributed Applications](https://sigops.org/s/conferences/sosp/2026/accepted.html) | 分散アプリケーションのソースレベル対話的デバッグ。実行時状態の横断観測 | 会議 | | 66 | IWQoS 2026 | [DelAct: A Replayable Boundary Runtime for Auditable and Governed LLM Agent Workflows](https://iwqos2026.ieee-iwqos.org/detailed-technical-sessions) | LLM エージェントのワークフローを再生可能にし監査と統制を可能にする実行時基盤 | 会議 | | 66 | 2026 IEEE Global Symposium on Emerging and Communication Technologies (GSEACT) | [Deterministic Lightweight Observability for Edge and Cloud-Native Applications](https://www.semanticscholar.org/paper/c25ce107156d5494ac3c410dba9d49b9aca42038) | エッジとクラウドネイティブ向けの決定的で軽量な可観測性。会場は小規模で内容未確認 | メール | | 66 | IWQoS 2026 | [DistSRE: Synthesizing Runtime Executions with Automated Co-Optimization for Distributed LLM Training](https://iwqos2026.ieee-iwqos.org/detailed-technical-sessions) | 分散 LLM 学習の実行時挙動を合成し自動的に共最適化 | 会議 | | 66 | ICWS 2026 | [Feedback Control-Based Self-Adaptive Cloud Resource Provisioning for Multi-tier Web Systems](https://services.conferences.computer.org/2026/icws-program/) | 多階層ウェブ系へのフィードバック制御に基づく自律的資源供給 | 会議 | | 66 | ICS 2026 | [GRASP: Fine-grained and Adaptive Sampled Simulation for GPU Performance Modeling](https://dipsa-qub.github.io/ICS2026-webpage/program/program.html) | 細粒度かつ適応的な標本抽出による GPU 性能模擬。サンプリングで測定コストを抑える | 会議 | | 66 | ICWS 2026 | [JPSTL: Joint Probabilistic Signal Temporal Logic for Predictive Composition Monitoring](https://services.conferences.computer.org/2026/icws-program/) | 確率的な信号時相論理で合成サービスの将来の違反を予測的に監視 | 会議 | | 66 | ICWS 2026 | [MIST: Microservice Identification from Sequential Traces](https://services.conferences.computer.org/2026/icws-program/) | 逐次トレースからマイクロサービスの構成単位を同定。トレースを入力とした構造推定 | 会議 | | 66 | ISCA 2026 | [Rearchitecting the Datacenter Lifecycle for AI](https://iscaconf.org/isca2026/program/) | AI 向けにデータセンターの一生を再設計する構想。クラウド事業者視点で追う価値 | 会議 | | 66 | ICDCS 2026 | [SLOShield: Contract-Bounded Safe Evolution for Interpretable Scheduling in Heterogeneous Accelerator Clusters](https://icdcs2026.icdcs.org/program/main-technical-sessions/) | 異種アクセラレータクラスタで SLO を契約として束縛し解釈可能なスケジューリングを進化させる | 会議 | | 66 | arXiv (Cornell University) | [TiRex-2: Generalizing TiRex to Multivariate Data and Streaming](https://www.semanticscholar.org/paper/34fa158b427c0dc5f2372b15376b8df31c28b91e) | ストリーミング下で一定コストを保つ再帰型の多変量時系列基盤モデル | メール | | 66 | ICWS 2026 | [V-Ideal: Generative Prompting for Fault-Tolerant Multimodal Web Services under Microservice Failures](https://services.conferences.computer.org/2026/icws-program/) | マイクロサービス障害下のウェブサービス耐障害性を生成的プロンプティングで支える | 会議 | | 65 | ICWS 2026 | [Ada-Jelly: A Lightweight and Adaptive Subflow Sketch for High-Speed Network Monitoring](https://services.conferences.computer.org/2026/icws-program/) | 高速ネットワーク監視向けの軽量かつ適応的なサブフロースケッチ | 会議 | | 65 | ICDCS 2026 | [Efficient Deployment of Lightweight LLMs on Edge Devices: Kernel-Level Profiling and Cross-Platform Insights](https://www.semanticscholar.org/paper/08320aa4e307c5849bdf7805bdfb6070915efe11) | エッジ端末上の軽量 LLM 推論をカーネルレベルで詳細プロファイリング | メール+会議 | | 65 | IWQoS 2026 | [HyperWeave: QoS-aware GPU Overcommitment for Deep Learning Training Jobs](https://iwqos2026.ieee-iwqos.org/detailed-technical-sessions) | 深層学習学習ジョブに対する QoS 考慮の GPU オーバーコミット | 会議 | | 65 | ISCA 2026 | [KernelEvolve: Scaling Agentic Kernel Coding for Heterogeneous AI Accelerators at Meta](https://www.semanticscholar.org/paper/f59553a2dd62d0ca379fe6463d0ae3d94c76e19c) | Meta 本番向けに異種アクセラレータのカーネル生成をエージェントで自動化した事例 | メール+会議 | | 65 | ICDCS 2026 | [PR-Reduce: Heterogeneity-Aware Proportional AllReduce for Distributed Training](https://icdcs2026.icdcs.org/program/main-technical-sessions/) | 異種構成クラスタで比例配分する AllReduce。ストラグラー緩和に直結 | 会議 | | 65 | IWQoS 2026 | [PSM: Timely and Resource-Efficient Sketch Migration in Network Measurement](https://iwqos2026.ieee-iwqos.org/detailed-technical-sessions) | ネットワーク計測用スケッチの移行を時間的にも資源的にも効率よく行う | 会議 | | 65 | ISCA 2026 | [PowerGrad: Hierarchical Power Management for Power-Limited ML Inference Clusters](https://iscaconf.org/isca2026/program/) | 電力制約下の推論クラスタを階層的に電力管理 | 会議 | | 65 | ICWS 2026 | [QoS-Aware Brokering for LLM-as-a-Service: SLA-Constrained Routing with Drift Detection](https://services.conferences.computer.org/2026/icws-program/) | LLM as a Service の SLA 制約付きルーティングにドリフト検知を組み込む | 会議 | | 65 | IWQoS 2026 | [Semantic Raft: A Fault-Tolerant Multi-Agent Inference Framework for Reliable LLM Services](https://iwqos2026.ieee-iwqos.org/detailed-technical-sessions) | 複数エージェントの合意で LLM サービスの障害耐性を高める枠組み | 会議 | | 65 | ASE 2026 | [Stop When It Matters: Detectability-Guided Microbenchmarking for Performance Regression Testing](https://conf.researchr.org/track/ase-2026/ase-2026-research-track) | 検出可能性を基準にマイクロベンチマークを打ち切る性能回帰検知。計測コストと検出力の折衷 | 会議 | | 65 | KDD 2026 | [TS-Memory: Plug-and-Play Memory for Time Series Foundation Models](https://kdd2026.kdd.org/papers/) | 時系列基盤モデルに着脱可能な記憶機構を付与する | 会議 | | 65 | SIGCOMM '26: ACM SIGCOMM 2026 Conference | [TurboBus: Pooling PCIe Bandwidth for LLM Workloads via Scale-Up Fabrics](https://www.semanticscholar.org/paper/2194e3289290c79fa67dbe7fed4b4ac336726b03) | スケールアップ用ファブリックで PCIe 帯域を融通。SIGCOMM の大規模実測が期待できる | メール | | 65 | ACM Transactions on Architecture and Code Optimization | [Xtream: A Production-Level VM Cross-Cloud Disk Migration System with Stripe-Oriented Prefetching](https://www.semanticscholar.org/paper/626f4ba99ba0d1bc0d41490c33c5acf326ecff19) | 実運用のクラウド間 VM ディスク移行システム。効果は数値で示すが自社報告値 | メール | | 64 | KDD 2026 | [Hierarchical Adversarial Bandits for Online Configuration Optimization](https://kdd2026.kdd.org/papers/) | 階層的な敵対的バンディットによるオンライン構成最適化 | 会議 | | 64 | ISSTA 2026 | [NSync: Automated Cloud Infrastructure-as-Code Reconciliation with AI Agents](https://conf.researchr.org/track/issta-2026/issta-2026-research-papers) | クラウドの IaC と実インフラの乖離を AI エージェントで調停。構成ドリフトの検知と是正 | 会議 | | 64 | IWQoS 2026 | [O3-Sketch: Memory-Efficient Online Chaotic Flow Detection in High-Speed Networks](https://iwqos2026.ieee-iwqos.org/detailed-technical-sessions) | 高速ネットワーク上でメモリ効率よく異常な流れを検知するスケッチ | 会議 | | 64 | ICDCS 2026 | [Prism: Proactive Workload-Aware Optimization for Hybrid-Service in LMaaS Systems](https://icdcs2026.icdcs.org/program/main-technical-sessions/) | 言語モデル提供基盤の混在ワークロードを予測して先回りで最適化 | 会議 | | 64 | ICS 2026 | [Skew-aware Adaptive All-to-allv Algorithms for Dynamic Deep Learning Workloads](https://dipsa-qub.github.io/ICS2026-webpage/program/program.html) | 動的な深層学習ワークロードの偏りを考慮した all-to-allv 集合通信 | 会議 | | 64 | IWQoS 2026 | [Speak Your Network: Automated Network Emulation Construction with Cost-Efficient Multi-Agent Orchestration](https://iwqos2026.ieee-iwqos.org/detailed-technical-sessions) | 自然言語からネットワークエミュレーション環境を多エージェントで構築。費用対効果も明示 | 会議 | | 63 | ICDCS 2026 | [BCD-Megatron: A Cost-Effective Training System for Large Language Models](https://www.semanticscholar.org/paper/655680833babd2d0a7fb300f8d417bc836a79809) | 大規模言語モデル学習のコスト効率を狙う学習システム | メール+会議 | | 63 | ICS 2026 | [CATS: Correlation-aware Task Scheduling for GPU Power Optimization in AI Data Centers](https://dipsa-qub.github.io/ICS2026-webpage/program/program.html) | 相関を手がかりにした GPU 電力最適化のためのタスクスケジューリング | 会議 | | 63 | ISCA 2026 | [Combating the Memory Walls: Optimization Pathways for Long-Context Agentic LLM Inference](https://iscaconf.org/isca2026/program/) | 長文脈エージェント推論のメモリ壁を対象とした最適化経路の整理 | 会議 | | 63 | ICS 2026 | [Latency-SLO-Aware Memory Offloading for Large Language Model Inference](https://dipsa-qub.github.io/ICS2026-webpage/program/program.html) | 遅延目標を意識したメモリオフロードによる LLM 推論制御。SLO を制御変数に据える | 会議 | | 63 | IWQoS 2026 | [Switch-Transparent Load Balancing for RDMA Data Centers: A Host-Only Approach](https://iwqos2026.ieee-iwqos.org/detailed-technical-sessions) | スイッチを変更せずホスト側のみで RDMA データセンターの負荷分散を実現 | 会議 | | 62 | IWQoS 2026 | [Biphasic Sketch: Multi-Attribute Stream Summarization for Arbitrary Attribute Combinations](https://iwqos2026.ieee-iwqos.org/detailed-technical-sessions) | 任意の属性組み合わせに対して多属性ストリームを要約するスケッチ | 会議 | | 62 | ASE 2026 | [CIPIHunter: Detecting Configuration-Induced Prediction Instability in Deep Learning Frameworks](https://conf.researchr.org/track/ase-2026/ase-2026-research-track) | 設定差異が推論結果を不安定化させる現象を検知。設定起因の不具合を系統的に捉える | 会議 | | 62 | IWQoS 2026 | [Chronus: Accurate and Memory-Efficient Sketching for High-Cumulative-RTT Flow Detection](https://iwqos2026.ieee-iwqos.org/detailed-technical-sessions) | 累積往復遅延の大きいフローをメモリ効率よく検出するスケッチ | 会議 | | 62 | KDD 2026 | [CloudCons: A Comprehensive End-to-End Benchmark for Cloud Resource Consolidation](https://kdd2026.kdd.org/papers/) | クラウド資源統合の端から端までのベンチマーク。評価設計を正面から扱う | 会議 | | 62 | IWQoS 2026 | [Dynamic Hot Expert Replication with Load and Topology-Aware Joint Gating for Distributed MoE Inference](https://iwqos2026.ieee-iwqos.org/detailed-technical-sessions) | 分散 MoE 推論における負荷とトポロジを考慮したエキスパート複製 | 会議 | | 62 | IEEE Transactions on Knowledge and Data Engineering | [EnsDiffAD: Ensemble Diffusion Models for Multivariate Time Series Anomaly Detection](https://www.semanticscholar.org/paper/7a8baaca934613047c763b428d7de580d5692ef3) | 拡散モデルの集団学習による多変量時系列異常検知。推論コストの重さが運用面の懸念 | メール | | 62 | SOSP 2026 | [Guiding Agentic GPU Kernel Optimization with Data Flow Invariants](https://sigops.org/s/conferences/sosp/2026/accepted.html) | データフロー不変量でエージェントによる GPU カーネル最適化を導く | 会議 | | 62 | ICDCS 2026 | [HiAsCC: Hierarchical Asynchronous Collective Communication Method for Large Model Training](https://icdcs2026.icdcs.org/program/main-technical-sessions/) | 大規模モデル学習向けの階層的非同期集合通信 | 会議 | | 62 | ISCA 2026 | [Mapping and Communication Optimizations with Fault Tolerance for Wafer-Scale LLM Inference](https://iscaconf.org/isca2026/program/) | ウェハスケール上の LLM 推論に耐障害性を組み込む。欠陥コアを前提とした設計 | 会議 | | 62 | IWQoS 2026 | [PRO: Deterministic Load Balancing for AI Datacenter Network](https://iwqos2026.ieee-iwqos.org/detailed-technical-sessions) | AI データセンターネットワークの決定的ロードバランシング | 会議 | | 62 | ISCA 2026 | [Power Sloshing in Compound Servers for Large-Scale AI Inference Workloads](https://www.semanticscholar.org/paper/f8d93ca9406f0e192a779e41603eca8ce48d6fab) | 大規模 AI 推論における複合サーバの電力変動 | メール+会議 | | 62 | KDD 2026 | [Rethinking Weak Supervision in Anomaly Detection: A Comprehensive Benchmark](https://kdd2026.kdd.org/papers/) | 弱教師あり異常検知を包括的に評価し直すベンチマーク | 会議 | | 62 | ICS 2026 | [SPPO: Making Million-Token LLM Training Practical on Modest GPU Clusters](https://dipsa-qub.github.io/ICS2026-webpage/program/program.html) | 控えめな GPU クラスタで百万トークン規模の LLM 学習を成立させる | 会議 | | 62 | ICS 2026 | [StreamGuard: Low-Overhead Resilience for Real-time HPC Data Streams](https://dipsa-qub.github.io/ICS2026-webpage/program/program.html) | 実時間データストリームに対する低オーバーヘッドの耐障害性。テレメトリ収集経路の信頼性設計と重なる | 会議 | | 62 | Lirias (KU Leuven) | [SubTSMD: discovering subspace motifs with temporal variations in multivariate time series](https://www.semanticscholar.org/paper/254cedbcdeafb8d8c1d25b05a2a6b1201b09e0c7) | 多変量時系列から部分空間モチーフを時間変動込みで発見。運用テレメトリへの転用余地 | メール | | 62 | SOSP 2026 | [Validating a Production Cloud Object Store with Lightweight Formal Methods](https://sigops.org/s/conferences/sosp/2026/accepted.html) | 実運用のクラウドオブジェクトストレージを軽量形式手法で検証した報告 | 会議 | | 60 | ASE 2026 | [Assessing Uncertainty in Performance Modeling: Conformal Prediction vs. Bayesian Regression](https://conf.researchr.org/track/ase-2026/ase-2026-research-track) | 性能モデルの不確実性を等角予測とベイズ回帰で比較評価。誤警報制御に転用の余地 | 会議 | | 60 | ASE 2026 | [Catching Developers in the Flow: Low-Latency Agentic Program Repair at Google Scale](https://conf.researchr.org/track/ase-2026/ase-2026-industry-showcase) | Google 規模での低遅延なエージェント型自動プログラム修復 | 会議 | | 60 | ASE 2026 | [ConfFuzz: Parameter-Aware Greybox Fuzzing for Configurable Cloud Systems](https://conf.researchr.org/track/ase-2026/ase-2026-research-track) | 設定パラメータを意識したグレイボックスファジング。設定起因の潜在的障害を掘り起こす | 会議 | | 60 | KDD 2026 | [Disentangling Multi-View Scanning in Mamba for Network Traffic Anomaly Detection](https://kdd2026.kdd.org/papers/) | Mamba の多視点走査を分離したネットワークトラフィック異常検知 | 会議 | | 60 | ISCA 2026 | [DynoPipe: Heterogeneous Edge-Cloud LLM Serving with Dynamically Orchestrated Pipeline Boundaries](https://www.semanticscholar.org/paper/cc4677b154afe001e2c151d565f6e5f81b927480) | エッジとクラウドをまたぐ推論パイプライン境界の動的再編 | メール+会議 | | 60 | ASE 2026 | [EffiHolmes: Differential Profiling-Guided Repository Level Time Inefficiency Fix Localization](https://conf.researchr.org/track/ase-2026/ase-2026-research-track) | 差分プロファイリングで時間非効率の修正箇所を特定 | 会議 | | 60 | ICWS 2026 | [Optimizing FaaS Platforms for MCP-enabled Agentic Workflows](https://services.conferences.computer.org/2026/icws-program/) | MCP を用いるエージェント型ワークフローに向けた FaaS 基盤の最適化 | 会議 | | 60 | KDD 2026 | [OrionInfer: Low-Overhead Parallelism Switching and Live Migration for Efficient LLM Serving](https://kdd2026.kdd.org/papers/) | LLM 推論の並列度切り替えと稼働中移送を低オーバーヘッドで行う | 会議 | | 60 | ICS 2026 | [Phase-aware Peak Power Reduction for Minimizing the Capital Expense of LLM Inference](https://dipsa-qub.github.io/ICS2026-webpage/program/program.html) | LLM 推論のフェーズを考慮したピーク電力削減で設備投資を抑える | 会議 | | 60 | IWQoS 2026 | [Pluto: Fast and Accurate Persistent Flow Detection in High-Speed Networks via Adaptive Protection](https://iwqos2026.ieee-iwqos.org/detailed-technical-sessions) | 高速ネットワークでの持続的フロー検知を適応的保護で高精度化 | 会議 | | 60 | ISCA 2026 | [PowerWeave: Unlocking Energy-Efficient ML on GPUs with OS-Level Spatial Power Management](https://iscaconf.org/isca2026/program/) | OS 層の空間的な電力管理により GPU 上の機械学習を省電力化 | 会議 | | 60 | ISCA 2026 | [Prometheus: Toward Resilient Datacenters through Optimized Cooling Infrastructure](https://iscaconf.org/isca2026/program/) | 冷却設備の最適化によるデータセンターの耐障害性向上 | 会議 | | 60 | IWQoS 2026 | [REACT: Toward Real-Time, End-to-End, Adaptive Cross-Layer Restoration for IP-Over-Optical Networks](https://iwqos2026.ieee-iwqos.org/detailed-technical-sessions) | IP 層と光層をまたぐ復旧を実時間かつ適応的に行う | 会議 | | 60 | ASE 2026 | [Understanding the Performance-Effectiveness Trade-Offs of AddressSanitizer in Real-World Programs: An Empirical Study](https://conf.researchr.org/track/ase-2026/ase-2026-industry-showcase) | 動的検査器の性能と有効性のトレードオフを実証的に測る。計装オーバーヘッドの代償を数値で示す | 会議 | ## 境界例 見送った候補と理由。50–59 のうち、profile のコア関心に近づきながら決め手を欠いたものを 8 件挙げる。 ### [Practical Software Performance Analysis on Linux Systems](https://www.semanticscholar.org/paper/d27a5e02fae65b2869ee2ef7400b449f3aa5a880) - 会場/年: FSE Companion '26: 34th ACM International Conference on the Foundations of Software Engineering / 2026 - スコア: 58 - 見送り理由: Linux 上の実践的な性能解析で動的計装に近いが、FSE の併設収録であり解説的発表とみられ新規知見の量が読めない。 ### [Self Healing and Policy Governed Cloud Native Systems for Long Running Distributed Applications](https://www.semanticscholar.org/paper/cb73a453184e6cfdbf845a388f87348d65e7b8b8) - 会場/年: 2026 IEEE Global Symposium on Emerging and Communication Technologies (GSEACT) / 2026 - スコア: 58 - 見送り理由: セルフヒーリングと方針統制という主題はコア関心に触れるが、abstract が無く会場の選抜性も低いため実証の規模を判断できない。 ### [Testing Custom Control Planes Without the Cluster](https://sigops.org/s/conferences/sosp/2026/accepted.html) - 会場/年: SOSP 2026 / 2026 - スコア: 58 - 見送り理由: 制御プレーンの誤動作は広範囲の障害を招くため予防価値はあるが、テレメトリや事後診断ではなく試験手法である。 ### [TSN-FLARE: A Flexible and Resilient Telemetry Framework for Time-Sensitive Networking](https://iwqos2026.ieee-iwqos.org/detailed-technical-sessions) - 会場/年: IWQoS 2026 / 2026 - スコア: 58 - 見送り理由: 柔軟かつ耐障害のテレメトリ基盤という主題は合致するが、対象が時間制約ネットワークでクラウド運用から離れる。 ### [ProbGuard: Proactive Runtime Monitoring for LLM Agent Safety via Probabilistic Prediction](https://conf.researchr.org/track/ase-2026/ase-2026-research-track) - 会場/年: ASE 2026 / 2026 - スコア: 58 - 見送り理由: 確率的予測による LLM エージェントの実行時監視だが、主眼が安全性の逸脱防止であり障害検知や原因特定とは目的が異なる。 ### [TTL Jumps: Unexpected TTL Rewrites Impacting Inferences from Traceroutes](https://conferences.sigcomm.org/imc/2026/accepted-papers/) - 会場/年: IMC 2026 / 2026 - スコア: 58 - 見送り理由: traceroute の推論を狂わせる TTL 書き換えを暴く計測研究で、計測手法の妥当性を問う姿勢は評価できるがサービス運用への直結は薄い。 ### [CacheWise: Understanding Workloads and Optimizing KVCache Management for Efficiently Serving LLM Coding Agents](https://www.semanticscholar.org/paper/1d58969ef96fb52ce1982f1816e06eb8d4c01703) - 会場/年: arXiv (Cornell University) / 2026 - スコア: 58 - 見送り理由: 実世界のコーディングエージェントのトレース特性解析は好ましいが、主眼は KV キャッシュ最適化であり arXiv のみで例外基準にも該当しない。 ### [RangeGuard: Efficient, Bounded Approximate Error Correction for Reliable DNNs](https://iscaconf.org/isca2026/program/) - 会場/年: ISCA 2026 / 2026 - スコア: 58 - 見送り理由: 無音のデータ破損への耐性という点で AI 基盤信頼性に触れるが、検知ではなく訂正が主眼でハードウェア設計に閉じる。 ## ingest コマンド 必読 44 件。会議 accepted papers はまだ個別ページを持たないものが多いため、URL が取れない候補はタイトルを渡す形式にした。 ``` /wiki-ingest-paper "GALA: Graph-Augmented LLM Agents for Root Cause Analysis and Incident Response in Microservices" /wiki-ingest-paper "LMTracer: Fine-Grained and Real-Time Performance Profiling for Production LLM Systems" /wiki-ingest-paper "RCAgentBench: An Agent-Oriented Benchmark for Multimodal Root Cause Analysis in Microservices" /wiki-ingest-paper https://www.semanticscholar.org/paper/c08684985da94b5c8448fa454c98140c87a8da4d /wiki-ingest-paper "Gleaner: A Semantically-Rich and Efficient Online Sampler for Microservice Diagnostics" /wiki-ingest-paper "QProf: Fleetwide Transitive Cost Profiling of Warehouse-Scale Services" /wiki-ingest-paper "QueueBreak: A Trace-to-Diagnosis Pipeline for Tail Latency in Agentic LLM Services" /wiki-ingest-paper "ActionNex: A Production-Grade System for Next-Best-Action Recommendation in Cloud Outage Management" /wiki-ingest-paper https://www.semanticscholar.org/paper/b1a5df7e0a4503209b89cb6a1bef6d106b179769 /wiki-ingest-paper "Enhancing Trace-Based Root Cause Analysis for Microservice Systems via Code Change Understanding" /wiki-ingest-paper "HERO: Hypothesis-Centered Root-Cause Analysis for Microservice Incidents" /wiki-ingest-paper "On-site, Non-speculative Failure Diagnosis with CLODS" /wiki-ingest-paper "OpsLens: Runbook-Free Root Cause Analysis via Multi-agent Collaboration and Tool Management" /wiki-ingest-paper "SCALER: LLM-based Cross-Modal Alignment for Microservice Root Cause Analysis" /wiki-ingest-paper "Smart Brain: Semantic Anomaly Detection for Operational Time Series in Large Scale Service Systems" /wiki-ingest-paper "CAVIAR: Disentangling Root Causes with an ICA-based VAE for Large-Scale Microservice Systems" /wiki-ingest-paper "ARMOR: A Robust Self-Supervised Framework for Root Cause Analysis in Microservices under Missing Modality" /wiki-ingest-paper "AlarmClaw: Context-Enriched Alarm Management with Category-/Severity-Aware Incident Graphs" /wiki-ingest-paper "DCRL: Dynamic Causal Representation Learning for Root Cause Localization in Microservice Systems" /wiki-ingest-paper https://www.semanticscholar.org/paper/a9588f5025be8f16bc1fb3c4e742999f260305fe /wiki-ingest-paper "Metra: A Fault Localization Technique for Microservice Systems via the Fusion of Metrics and Traces" /wiki-ingest-paper "TD-RCA: Topology-Aware and Dual-Perspective Decoupling for Root Cause Analysis in Microservices" /wiki-ingest-paper "Towards Reliable Alarm Flood Reduction in Data Center via Grammar-Guided Verifiable LLM Reasoning" /wiki-ingest-paper "TuxBot: Semantic-Aware Online OS Tuning with LLMs" /wiki-ingest-paper "Waiting at the front door: Continuous monitoring of latency in the host network stack" /wiki-ingest-paper "When Does AI Actually Help in Incident Response? Identifying Good First Messages in Cloud Service Incidents" /wiki-ingest-paper "DualLane: Fast and Reliable LLM Agents for Interactive AIOps via Dual-Path Planning" /wiki-ingest-paper "LMID: A Comprehensive Multimodal Dataset for Failure Prediction in Cloud Computing Systems" /wiki-ingest-paper "CAPMix: Robust KPI Anomaly Detection for AIOps in Noisy and Dynamic Environments" /wiki-ingest-paper "Few-shot Multimodal Anomaly Detection via Dynamic Intra-modal Sparsity Attention and Quality-aware Cross-modal Fusion in Microservice System" /wiki-ingest-paper "LogNexus: Effective Log Compression via Unified Redundancy Encoding" /wiki-ingest-paper "Ranking Anomalous Subsequences in Large Scale Streaming Data for Efficient Triage" /wiki-ingest-paper "Tracing the Invisible: Semantic Message Flow Discovery via Data Contracts in Real-World Distributed Systems" /wiki-ingest-paper https://www.semanticscholar.org/paper/f6ed1160493df8dc3ba65232c342afcfd06a1f82 /wiki-ingest-paper "Discovering Performance Archetypes: Critical-Path-Aware Pattern Analysis and Regression Detection" /wiki-ingest-paper "Falconf: Configuration Error Diagnosis via Log Sequence Learning and Automated Misconfiguration Injections" /wiki-ingest-paper "In a Streaming World, Should You Stand Still? A Comprehensive Benchmark of Anomaly Detection in Streams" /wiki-ingest-paper "Mantis: Decoding HPC Telemetry Data for Robust System Prediction" /wiki-ingest-paper https://www.semanticscholar.org/paper/bd792779baf8a4167f30c6b82c854b99dc8daa39 /wiki-ingest-paper "SG-LR QoS: Sparse-Gating and Layered Low-Rank Based Lightweight QoS Monitoring for Multi-tenant Clouds" /wiki-ingest-paper "StepFinder: A Temporal Semantic Framework for Failure Attribution in Multi-Agent Systems" /wiki-ingest-paper "TORCL: Lightweight Root Cause Localization with Temporal-Order-Aware Ranking in Microservices" /wiki-ingest-paper "Taming Multi-Dimensional Tail: A Factorized Tail-Interaction Framework for Hyperscale Cloud Workload Forecasting" /wiki-ingest-paper "TuneAgent: Agentic Operating System Kernel Tuning with Reinforcement Learning" ``` 注目 138 件の ingest は上表のリンク列から個別に選ぶ。 ## 収集統計 - 収集元別件数: - Semantic Scholar 推薦メール: 89 件抽出(週次ダイジェスト 1 通、遡及 7 日) - 会議プログラム巡回: 10 会議回・376 件抽出 | 会議回 | 抽出方針 | 絞り込み前 | 抽出 | |---|---|---|---| | SOSP-2026 | full | 62 | 62 | | IMC-2026 | full | 23 | 23 | | ICDCS-2026 | full | 136 | 30 | | IWQoS-2026 | full | 128 | 48 | | ICWS-2026 | full | 135 | 35 | | ISSTA-2026 | full | 210 | 11 | | ASE-2026 | full | 338 | 51 | | ISCA-2026 | filtered | 173 | 43 | | ICS-2026 | filtered | 97 | 33 | | KDD-2026 | filtered | 1415 | 40 | - 巡回対象外: 既チェック 11 会議回(DSN・ICPE・WWW・OSDI・VLDB・SIGMOD・SIGCOMM・INFOCOM・ESEC/FSE・SSE・CLOUD の 2026 回)、開催窓外 27 会議回、採択未公開 3 会議回(APSys・IC2E・PRDC)、産業カンファレンス 5 件(SREcon ほか) - 巡回失敗: ICML-2026(静的 JSON が先頭 200 件のみ・API が 403)、SOSE-2026(プログラムが PDF のみ)。いずれも `checked_conferences` に記録せず翌週再試行する - 除外件数: - メタデータ照会失敗: 7 件 - `wiki/sources/` に既存: 21 件(メール側 6・会議側 15) - 提示済み(`.state.json`): 0 件 - 両収集器で重複を統合: 11 件 - 低スコア(60 未満)で見送り: 244 件(うち 50–59 が 92 件、50 未満が 152 件) - 候補総数(重複排除後): 426 件 / 提示: 190 件(必読 44・注目 138・境界例 8) - ③ 埋め込み一次絞り込み: `strategy: embedding`(ollama + nomic-embed-text、`wiki/sources/` 1369 ページを索引化) ### 今回の運用上の逸脱 - **Semantic Scholar API を使えなかった**。1Password 内の API キーが 403(失効の可能性)、keyless は 429 で恒常的に拒否された。ユーザー承認のうえ **OpenAlex を主経路、Crossref と arXiv を補完**としてメタデータを取得した。外部へ送ったのは論文タイトルと識別子のみである。 - **abstract がほとんど取得できなかった**。2026 年の accepted papers は要旨未公開のものが多く、426 候補のうち abstract が付いたのは 37 件にとどまる。 - **③ の上位 40 件による足切りを行わず、426 件すべてを ④ の LLM 再判定へ回した**。abstract 欠落によりタイトルのみの埋め込みとなり、一次スコアが 0.81–0.69 に圧縮されて識別力を失ったためである。実際、必読(80+)に入った 44 件のうち 8 件は一次順位が 200 位以下だった(Gleaner 218 位・CLODS 305 位・LogNexus 330 位・IMC の遅延監視 365 位・Tracing the Invisible 376 位・QProf 396 位ほか)。上位 40 件で切っていればこれらを取りこぼしていた。 - **④ の判定は abstract 抜きのタイトルと会場に依拠する部分が大きい**。各候補の推薦理由には確信度の低さを明記させたが、通常週より判定の精度は落ちる。ingest 前に原文の確認を推奨する。 - **注目(60–79)を一覧表形式にした**。138 件と多く、雛形どおりの節形式では実用に耐えないため。全件を落とさず掲載している。 - 会議収集器は応答長の制約から、`full` 指定の ICDCS・IWQoS・ICWS・ISSTA・ASE についても全件取得後にタイトル段階の意味的スクリーニングを適用した(上表の「絞り込み前 → 抽出」)。`full` の建前は取りこぼさないことなので、この点は次回の改善対象とする。