# Shengkun Cui [[@2025__SC__Characterizing GPU Resilience and Impact on AI - HPC Systems]](Cui+, SC2025、別題 "Story of Two GPUs")の共同筆頭著者(Archit Patke と equal contribution)。所属は [[University of Illinois Urbana-Champaign]](Urbana, USA、[email protected])。同論文では [[NCSA]] の [[Delta]] における A100/H100 GPU の 2.5 年分のレジリエンス特徴付けを主導した。 ## 関連 - ソース: [[@2025__SC__Characterizing GPU Resilience and Impact on AI - HPC Systems]] - 組織: [[University of Illinois Urbana-Champaign]] / [[NCSA]] - 人物: [[Ravishankar K. Iyer]](責任著者) - 概念: [[GPUクラスタ運用]] - [[Kaleidoscope]] フレームワーク([[@2020__SC20__Live Forensics for HPC Systems - A Case Study on Distributed Storage Systems]], SC 2020)の共著者。[[Saurabh Jha]] とともに [[Blue Waters]] ストレージの障害フォレンジクスに取り組んだ。(Source: [[@2020__SC20__Live Forensics for HPC Systems - A Case Study on Distributed Storage Systems]]) - [[@2026__DSN__PRAXIS - Integrating Program Analysis with Observability for Root-Cause Analysis]](Cui+, DSN 2026)の筆頭著者。[[Ravishankar K. Iyer]]と共に[[University of Illinois Urbana-Champaign]]所属([email protected])で、[[IBM Research]]の[[Rahul Krishna]]・[[Saurabh Jha]]と共著。サービス依存グラフ(SDG)とプログラム依存グラフ(PDG)へのLLM駆動グラフトラバーサルによるクラウドRCAエージェントPRAXISを提案し、ReActベースラインに対しRCA精度6.3倍・トークン消費5.3倍削減を報告した。GPUレジリエンス研究(HPC/ハードウェア障害特徴付け)から、コード・設定起因のクラウドソフトウェア障害診断へと研究領域を広げている。(Source: [[@2026__DSN__PRAXIS - Integrating Program Analysis with Observability for Root-Cause Analysis]]) ## 出典 - [[@2025__SC__Characterizing GPU Resilience and Impact on AI - HPC Systems]](筆頭著者として登場) - [[@2020__SC20__Live Forensics for HPC Systems - A Case Study on Distributed Storage Systems]](共著者として登場) - [[@2026__DSN__PRAXIS - Integrating Program Analysis with Observability for Root-Cause Analysis]](筆頭著者として登場)