# How to Scale Your Model ## 概要 Jacob Austin, Sholto Douglas, Roy Frostig, Anselm Levskaya, Charlie Chen, Sharad Vikram, Federico Lebron, Peter Choy, Vinay Ramasesh, Albert Webson, Reiner Pope(執筆時点で一部は MatX 所属)による、Google DeepMind 発の Transformer スケーリング解説シリーズ。TPU/GPU 上での学習・推論における FLOPs・メモリ・通信の roofline 分析を軸に、システムレベルのシャーディング戦略までを一貫して扱う。全 9 部構成のオンライン書籍で、[jax-ml.github.io/scaling-book](https://jax-ml.github.io/scaling-book) で公開されている。(Source: [[@2025__ScalingBook__How to Scale Your Model - Part 7 Inference]]) ## 構成と主要テーマ(取り込んだ章) → [[@2025__ScalingBook__How to Scale Your Model - Part 7 Inference]] — Transformer 推論(prefill/generation)の roofline 分析、[[KVキャッシュ]]、[[Grouped-Query Attention|GQA]]・[[PagedAttention]]・[[Speculative Decoding|投機的サンプリング]]等の高速化技法、[[Prefill-Decode分離]]と[[動的バッチングと継続的バッチング|continuous batching]]による推論エンジン設計を扱う第 7 部 ## シリーズ構成(参考、本 wiki 未取り込み) Part 7 本文が言及する構成から、本書は Part 1(Roofline)〜Part 9(JAX 実装)の全 9 部からなると分かる。Part 7 が参照する関連部: - Part 4: Transformer の数理(FLOPs/メモリのモデル化) - Part 5: 学習のシャーディング戦略 - Part 6: LLaMA の学習への応用 - Part 8: LLaMA の推論への応用(Part 7 の続き) (Source: [[@2025__ScalingBook__How to Scale Your Model - Part 7 Inference]]) ## 引用 ``` Austin et al., "How to Scale Your Model", Google DeepMind, online, 2025. ``` ## 関連 - エンティティ: [[wiki/entities/Google TPU|Google TPU]] — 分析対象ハードウェアの中心 - 概念: [[KVキャッシュ]] / [[演算強度]] ## 出典 - [[@2025__ScalingBook__How to Scale Your Model - Part 7 Inference]](書誌情報、著者陣、シリーズ構成への言及)