# ByteDance Seed
[[ByteDance]] の研究組織。集団通信の依存関係トレーシング [[Mycroft]]([[@2025__SOSP__Mycroft - Tracing Dependencies in Collective Communication Towards Reliable LLM Training]])の複数著者の所属。(Source: [[@2025__SOSP__Mycroft - Tracing Dependencies in Collective Communication Towards Reliable LLM Training]])
大規模 LLM RL システム DAPO([[@2025__arXiv__DAPO - An Open-Source LLM Reinforcement Learning System at Scale]])の主著機関でもあり、プロジェクトリード [[Qiying Yu]] を含む Algorithm/Infrastructure/Dataset の各チームが所属する(Source: [[@2025__arXiv__DAPO - An Open-Source LLM Reinforcement Learning System at Scale]])。
FSDP システム veScale-FSDP([[@2026__MLSys2026__veScale-FSDP - Flexible and High-Performance FSDP at Scale]])の主著機関でもあり、[[Zezhou Wang]]・[[Yanghua Peng]]・[[Xin Liu]] らが所属する。veScale-FSDP は 10K GPU 超の本番訓練ワークロードにすでに展開されている。(Source: [[@2026__MLSys2026__veScale-FSDP - Flexible and High-Performance FSDP at Scale]])
マルチモーダルLLM訓練システム MegaScale-Omni([[@2026__EuroSys__MegaScale-Omni - A Hyper-Scale, Workload-Resilient System for MultiModal LLM Training in Production]]、EuroSys 2026)の主著機関でもあり、[[Yanghua Peng]]・[[Xin Liu]] を含む複数著者が所属する。encoderとLLMバックボーンの並列化を分離しつつ同一GPU集合上でコロケーションする encoder-LLM multiplexing により、数千GPU規模の社内MLLM訓練基盤として展開されている。(Source: [[@2026__EuroSys__MegaScale-Omni - A Hyper-Scale, Workload-Resilient System for MultiModal LLM Training in Production]])
GPU の Silent Data Corruption(SDC)を扱う 2 本の OSDI 2026 論文にも中心的に関与する。[[Shanghai Jiao Tong University]] との共同研究である SDCHunter([[@2026__OSDI__SDCs in the Wild - Characterizing and Diagnosing SDC-defective GPUs in Production LLM Training]])は本番クラスタから回収した SDC 欠陥 GPU 23 台を特性調査し決定論的リプレイによる診断システムを提案し、[[Tsinghua University]] との共同研究である AEGIS([[@2026__OSDI__Safeguarding LLM Training at Scale - Online SDC Detection and Insights from 35 Million GPU Hours]])はオンラインで SDC をリアルタイム検知する。両論文は共著者(Yun Zhang・Gaohong Liu・Zuquan Song・Shuguang Wang・[[Wencong Xiao]]・[[Xin Liu]] ら)が重複し、ByteDance の本番環境ではオンライン検知(AEGIS)とオフライン診断・局在化(SDCHunter)が相補的に連携する(SDCHunter 論文 §8 は AEGIS を SDC 検知ツールとして明示的に参照する)。(Source: [[@2026__OSDI__SDCs in the Wild - Characterizing and Diagnosing SDC-defective GPUs in Production LLM Training]], [[@2026__OSDI__Safeguarding LLM Training at Scale - Online SDC Detection and Insights from 35 Million GPU Hours]])
## 関連
- ソース: [[@2025__SOSP__Mycroft - Tracing Dependencies in Collective Communication Towards Reliable LLM Training]] / [[@2025__arXiv__DAPO - An Open-Source LLM Reinforcement Learning System at Scale]] / [[@2026__MLSys2026__veScale-FSDP - Flexible and High-Performance FSDP at Scale]] / [[@2026__EuroSys__MegaScale-Omni - A Hyper-Scale, Workload-Resilient System for MultiModal LLM Training in Production]] / [[@2026__OSDI__SDCs in the Wild - Characterizing and Diagnosing SDC-defective GPUs in Production LLM Training]] / [[@2026__OSDI__Safeguarding LLM Training at Scale - Online SDC Detection and Insights from 35 Million GPU Hours]]
- 親組織: [[ByteDance]]
- 関連プロダクト: [[Mycroft]] / [[veScale]] / SDCHunter / AEGIS
- 関連研究者: [[Qiying Yu]] / [[Zezhou Wang]] / [[Yanghua Peng]] / [[Wencong Xiao]] / [[Xin Liu]]
- 協力機関: [[Shanghai Jiao Tong University]] / [[Tsinghua University]]