# NVIDIA Hopper NVIDIA の GPU アーキテクチャ。H100 を中心に、FP8 対応 Tensor Core、Transformer Engine、thread block cluster、DPX、Tensor Memory Accelerator、分散共有メモリ、第4世代 NVLink を導入した。 公式発表は H100 SXM5 の132 SM・80GB HBM3・3TB/s超の帯域、AI 学習最大9倍・LLM 推論最大30倍などを示すが、これらは比較条件付きの出荷前推定値である([[@2022__NVIDIA Developer Blog__NVIDIA Hopper Architecture In-Depth]])。 後続の H800 マイクロベンチマークは、`wgmma`、FP8、TMA、分散共有メモリの実効性能が行列形状、クラスタサイズ、形式変換、メモリ律速性に依存することを示した([[@2024__arXiv__Benchmarking and Dissecting the Nvidia Hopper GPU Architecture]])。 GB203/RTX 5080 との比較では、GH100/H100 PCIe が FP64、HBM2e の global memory 帯域、短い依存命令列、大規模 FP8 D-GEMMで有利な条件を示した。(Source: [[@2025__arXiv__Dissecting the NVIDIA Blackwell Architecture with Microbenchmarks]]) ## 関連 - [[NVIDIA]] - [[@2024__arXiv__Benchmarking and Dissecting the Nvidia Hopper GPU Architecture]] - [[@2025__arXiv__Dissecting the NVIDIA Blackwell Architecture with Microbenchmarks]]