# Harvard University 米国の研究大学。[[Minder]] 論文([[@2025__NSDI__Minder - Faulty Machine Detection for Large-scale Distributed Model Training]])の責任著者 [[Minlan Yu]] の所属。(Source: [[@2025__NSDI__Minder - Faulty Machine Detection for Large-scale Distributed Model Training]]) 2019年当時は [[Yuliang Li]]・[[Minlan Yu]] の所属として [[HPCC]] 論文([[@2019__SIGCOMM__HPCC - High Precision Congestion Control]]、SIGCOMM 2019)にも参加し、[[Alibaba Group]] 主導の INT ベース高精度輻輳制御の開発に協力した。(Source: [[@2019__SIGCOMM__HPCC - High Precision Congestion Control]]) [[Astral]]([[@2025__SIGCOMM__Astral - A Datacenter Infrastructure for Large Language Model Training at Scale]], SIGCOMM 2025)の共著者 [[ChonLam Lao]] の所属でもある。(Source: [[@2025__SIGCOMM__Astral - A Datacenter Infrastructure for Large Language Model Training at Scale]]) [[Minlan Yu]] は集団通信の信頼性研究にも関与し、[[Mycroft]]([[@2025__SOSP__Mycroft - Tracing Dependencies in Collective Communication Towards Reliable LLM Training]], SOSP 2025)と 10万+GPU 規模の集団通信([[@2025__arXiv__Collective Communication for 100k+ GPUs]])の共著者である。(Source: [[@2025__SOSP__Mycroft - Tracing Dependencies in Collective Communication Towards Reliable LLM Training]], [[@2025__arXiv__Collective Communication for 100k+ GPUs]]) [[Minlan Yu]]・[[ChonLam Lao]]・[[Brian Sutioso]] は [[PrismLLM]]([[@2026__arXiv__A Few GPUs, A Whole Lotta Scale]])の共著者として、[[Alibaba Group]] 主導の少数 GPU による大規模 LLM 訓練エミュレーションシステムに参加した。(Source: [[@2026__arXiv__A Few GPUs, A Whole Lotta Scale]]) 時系列分野では、[[UniTS]]([[@2024__NeurIPS__UniTS - A Unified Multi-Task Time Series Model]]、NeurIPS 2024)の著者Shanghua Gao・Owen Queen・Marinka Zitnik(責任著者)が Harvard University に所属し、[[MIT Lincoln Laboratory]]・[[University of Virginia]]と共同で統一マルチタスク時系列モデルを開発した。(Source: [[@2024__NeurIPS__UniTS - A Unified Multi-Task Time Series Model]]) 2005年当時は [[Alexandra Fedorova]] の所属でもあり、同氏は [[@2005__HotOS__The Many Faces of Systems Research - And How to Evaluate Them]](HotOS 2005)の共著者としてシステム研究の評価方法論(科学・工学・芸術の3次元)を論じた。(Source: [[@2005__HotOS__The Many Faces of Systems Research - And How to Evaluate Them]]) ## 関連 - ソース: [[@2025__NSDI__Minder - Faulty Machine Detection for Large-scale Distributed Model Training]] / [[@2025__SIGCOMM__Astral - A Datacenter Infrastructure for Large Language Model Training at Scale]] / [[@2025__SOSP__Mycroft - Tracing Dependencies in Collective Communication Towards Reliable LLM Training]] / [[@2025__arXiv__Collective Communication for 100k+ GPUs]] / [[@2026__arXiv__A Few GPUs, A Whole Lotta Scale]] / [[@2024__NeurIPS__UniTS - A Unified Multi-Task Time Series Model]] / [[@2005__HotOS__The Many Faces of Systems Research - And How to Evaluate Them]] / [[@2019__SIGCOMM__HPCC - High Precision Congestion Control]] / [[@2025__JMLR__Scaling Data-Constrained Language Models]] - エンティティ: [[Minlan Yu]] / [[ChonLam Lao]] / [[Brian Sutioso]] / [[UniTS]] / [[MIT Lincoln Laboratory]] / [[University of Virginia]] / [[Alexandra Fedorova]] / [[Yuliang Li]] / [[Alibaba Group]] / [[Boaz Barak]]