SA Article Coverage Review · 2026-02-16_inferencex-v2-nvidia-blackwell-vs

Coverage Summary

  • Source: 開啟原始 SA 文章
  • Atomic Claims: 574
  • Source blocks: 410
  • Blocks with ≥1 Atomic Claim: 166
  • Blocks without Atomic Claim: 244
  • Unplaced Claims: 0

Coverage Review

請從頭到尾閱讀下方 SA 全文。Atomic Claim 會依 Evidence 在原文出現的位置 inline 插入。
若該英文段落已有繁中翻譯 cache,翻譯只會作為淡色閱讀輔助顯示;不會進入 source、Claim provenance 或 Graph。
沒有 Claim callout 的段落不一定有問題;若內容重要且應形成知識,請記到 Missing Claim Notes。

Missing Claim Notes

  • 若看到重要但沒有 Atomic Claim 的段落,請在這裡記錄:
    • Section:
    • Evidence:
    • 為什麼重要/應該抽成什麼 Claim:

SA Full Text + Translation + Atomic Claims

InferenceX v2: NVIDIA Blackwell Vs AMD vs Hopper - Formerly InferenceMAX

Introduction

InferenceXv2 (formerly InferenceMAX) builds on the foundation established by InferenceMAXv1, our open-source, continuously updated inference benchmark that has set a new standard for AI inference performance and economics. InferenceMAXv1 moved beyond static, point-in-time benchmarks by running continuous tests across hundreds of chips and popular open-source frameworks. Free dashboard available here.

InferenceXv2(前身為 InferenceMAX)建立在 InferenceMAXv1 的基礎之上。InferenceMAXv1 是我們開源、持續更新的 inference benchmark ↗,為 AI inference performance 與 economics 建立了新的標準。它不再只是靜態、單一時間點的 benchmark,而是持續在數百顆 chip 與主流 open-source framework 上反覆測試。免費 dashboard 可在此查看 ↗。

Atomic Claim 1/574 · 2026-02-16_inferencex-v2-nvidia-blackwell-vs::IX2-0001

Claim: InferenceXv2(前稱 InferenceMAX)建立在 InferenceMAXv1 的基礎上;後者是一套開源、持續更新的推論 benchmark,為 AI inference 的效能與經濟性建立了新的標準。
Frame: ATTRIBUTE · Mode: ASSERTED · Mapping: PARTIAL
開啟逐條審核

Atomic Claim 2/574 · 2026-02-16_inferencex-v2-nvidia-blackwell-vs::IX2-0002

Claim: InferenceMAXv1 不再只做單一時間點的靜態 benchmark,而是持續測試數百款晶片與主流開源框架。
Frame: ATTRIBUTE · Mode: ASSERTED · Mapping: COMPLETE
開啟逐條審核

Our benchmark has been widely reproduced, validated and/or supported by almost every major buyer of compute from Google Cloud to Microsoft Azure to Oracle, OpenAI , and many more.

我們的 benchmark 已被幾乎所有主要 compute 買家廣泛重現、驗證或支援 ↗,包括 Google Cloud ↗、Microsoft Azure ↗、Oracle、OpenAI ↗ 等許多業者。

Atomic Claim 3/574 · 2026-02-16_inferencex-v2-nvidia-blackwell-vs::IX2-0003

Claim: 這套 benchmark 已被幾乎所有主要算力買方廣泛重現、驗證或支持,包括 Google Cloud、Microsoft AzureOracleOpenAI 等。
Frame: ATTRIBUTE · Mode: ASSERTED · Mapping: PARTIAL
開啟逐條審核

InferenceXv2 builds on this foundation. It expands coverage to include large scale DeepSeek MoE disaggregated inference (disagg prefill, or simply “disagg”) with wide expert parallelism (wideEP) optimization to **all 6 NVIDIA western GPU SKUs from the past 4 years **as well as to every single AMD western GPU SKU released in the past 3 years – in total InferenceXv2 utilizes close to 1000 frontier GPUs for a full benchmark run across all SKUs.

InferenceXv2 在這個基礎上進一步擴大範圍,加入 large-scale DeepSeek MoE disaggregated inference(disagg prefill,以下簡稱 disagg)與 wide expert parallelism(wideEP)最佳化,涵蓋過去四年所有 6 款 NVIDIA 西方市場 GPU SKU,以及過去三年每一款 AMD 西方市場 GPU SKU。完整跑一輪所有 SKU,InferenceXv2 總共會動用接近 1,000 顆 frontier GPU。

Atomic Claim 4/574 · 2026-02-16_inferencex-v2-nvidia-blackwell-vs::IX2-0006

Claim: InferenceXv2 涵蓋過去四年間 NVIDIA 在西方市場推出的全部 6 款 GPU SKUs。
Frame: NARY_RELATION · Mode: ASSERTED · Mapping: COMPLETE
開啟逐條審核

Atomic Claim 5/574 · 2026-02-16_inferencex-v2-nvidia-blackwell-vs::IX2-0004

Claim: InferenceXv2 延續並擴展這套基礎。
Frame: ATTRIBUTE · Mode: ASSERTED · Mapping: PARTIAL
開啟逐條審核

Atomic Claim 6/574 · 2026-02-16_inferencex-v2-nvidia-blackwell-vs::IX2-0007

Claim: InferenceXv2 涵蓋過去三年間 AMD 在西方市場推出的所有 GPU SKUs。
Frame: NARY_RELATION · Mode: ASSERTED · Mapping: COMPLETE
開啟逐條審核

Atomic Claim 7/574 · 2026-02-16_inferencex-v2-nvidia-blackwell-vs::IX2-0005

Claim: InferenceXv2 涵蓋大規模 DeepSeek MoE disaggregated inference,並搭配 wide expert parallelism
Frame: ATTRIBUTE · Mode: ASSERTED · Mapping: PARTIAL
開啟逐條審核

Atomic Claim 8/574 · 2026-02-16_inferencex-v2-nvidia-blackwell-vs::IX2-0008

Claim: 一次完整的 InferenceXv2 benchmark run 會使用接近 1,000 顆 frontier GPUs
Frame: RELATION · Mode: ASSERTED · Mapping: COMPLETE
開啟逐條審核

With today’s release, InferenceXv2 is now the first suite to benchmark the Blackwell Ultra GB300 NVL72 and B300 across the whole pareto frontier curve, and it is the first third party benchmark to test disagg+wideEP multi-node FP4 and FP8 MI355X performance. In future iterations of InferenceX, we will continue to focus heavily on disaggregated serving with wide expert parallelism as that is what is deployed in production at Frontier AI Labs like OpenAI, Anthropic, xAI, Google Deepmind, DeepSeek as well as advanced API providers like TogetherAI, Baseten, and Fireworks. In this article, we will also break down the system engineering principles and economics in play around the latest Claude Code Fast mode feature .

隨著今天版本發布,InferenceXv2 成為第一套沿完整 Pareto frontier curve 測試 Blackwell Ultra GB300 NVL72 與 B300 的 benchmark suite,同時也是第一個第三方 benchmark,測試 MI355X 在 disagg+wideEP multi-node FP4/FP8 下的 performance。未來 InferenceX 仍會把 disaggregated serving + wide expert parallelism 當成重點,因為 OpenAI、Anthropic、xAI、Google DeepMind、DeepSeek 等 Frontier AI Lab,以及 TogetherAI、Baseten、Fireworks 等先進 API provider,production 環境實際就是這樣部署。本篇也會拆解最新 Claude Code Fast mode ↗ 背後的 system engineering 原則與 economics。

Atomic Claim 9/574 · 2026-02-16_inferencex-v2-nvidia-blackwell-vs::IX2-0009

Claim: InferenceXv2 是第一套針對 GB300 NVL72B300 進行完整 Pareto frontier curve benchmark 的測試套件。
Frame: ATTRIBUTE · Mode: ASSERTED · Mapping: COMPLETE
開啟逐條審核

Atomic Claim 10/574 · 2026-02-16_inferencex-v2-nvidia-blackwell-vs::IX2-0010

Claim: InferenceXv2 是第一套第三方 benchmark,可測試 multi-node MI355XFP4FP8 下的 disagg+wideEP 效能。
Frame: ATTRIBUTE · Mode: ASSERTED · Mapping: COMPLETE
開啟逐條審核

Atomic Claim 11/574 · 2026-02-16_inferencex-v2-nvidia-blackwell-vs::IX2-0011

Claim: 未來版本的 InferenceX 將持續高度聚焦搭配 wide expert parallelismdisaggregated serving
Frame: NARY_RELATION · Mode: EXPECTED · Mapping: COMPLETE
開啟逐條審核

Atomic Claim 12/574 · 2026-02-16_inferencex-v2-nvidia-blackwell-vs::IX2-0011::SPLIT02

Claim: OpenAIAnthropicxAIGoogle DeepmindDeepSeek 等 Frontier AI Labs,以及 TogetherAI、Baseten、Fireworks 等 API providers,在 production 中部署搭配 wide expert parallelismdisaggregated serving
Frame: NARY_RELATION · Mode: ASSERTED · Mapping: PARTIAL
開啟逐條審核

Our benchmark is completely open-source under Apache 2.0 – this means that we are able to move at the same rapid speed at which the AI software ecosystem is advancing. If you like our work and would like to show us some support, please drop a star on our GitHub ! We also provide a free data visualizer at https://inferencex.com for everyone in the ML community to explore the complete dataset themselves.

我們的 benchmark 完全以 Apache 2.0 open source,這讓我們能跟 AI software ecosystem 一樣快速迭代。如果你喜歡這項工作,也歡迎到 GitHub 幫我們按個 star ↗。我們另外提供免費 data visualizer(inferencex.com)↗,讓整個 ML community 都可以自行探索完整 dataset。

Atomic Claim 13/574 · 2026-02-16_inferencex-v2-nvidia-blackwell-vs::IX2-0013

Claim: inferencex 也提供免費 data visualizer,讓 ML 社群可自行探索完整 dataset。
Frame: ATTRIBUTE · Mode: ASSERTED · Mapping: PARTIAL
開啟逐條審核

Atomic Claim 14/574 · 2026-02-16_inferencex-v2-nvidia-blackwell-vs::IX2-0012

Claim: 關於 InferenceX:這套 benchmark 完全以 Apache 2.0 開源,因此能以接近 AI software ecosystem 演進的速度快速更新。
Frame: ATTRIBUTE · Mode: ASSERTED · Mapping: PARTIAL
開啟逐條審核

We will add DeepSeekv4 and other popular Chinese frontier models with day 0 support as over the past 6 months, we now have cleaned up a lot of tech debt and are able to move fast with stable infrastructure . We will also be adding TPUv7 Ironwood and Trainium3 to InferenceX later this year! If you want to contribute to our impactful mission while earning a competitive compensation, consider applying here .

接下來會把 DeepSeekv4 與其他熱門中國 frontier model 納入 Day 0 support。過去六個月我們清掉了大量 tech debt,現在 infrastructure 已經穩定,可以更快往前推 ↗。今年稍晚也會把 TPUv7 Ironwood 與 Trainium3 加入 InferenceX。若你也想參與這項有影響力的工作,同時取得具競爭力的薪酬,也可以考慮申請加入 ↗。

Atomic Claim 15/574 · 2026-02-16_inferencex-v2-nvidia-blackwell-vs::IX2-0014

Claim: InferenceX 計畫對 DeepSeekv4 與其他熱門中國 frontier models 提供 day-0 support。
Frame: ATTRIBUTE · Mode: ASSERTED · Mapping: COMPLETE
開啟逐條審核

Atomic Claim 16/574 · 2026-02-16_inferencex-v2-nvidia-blackwell-vs::IX2-0015

Claim: InferenceX 團隊表示,過去六個月已清理大量 technical debt。
Frame: ATTRIBUTE · Mode: ASSERTED · Mapping: COMPLETE
開啟逐條審核

Atomic Claim 17/574 · 2026-02-16_inferencex-v2-nvidia-blackwell-vs::IX2-0016

Claim: InferenceX 團隊表示,目前基礎設施已足夠穩定,可以更快速迭代。
Frame: ATTRIBUTE · Mode: ASSERTED · Mapping: COMPLETE
開啟逐條審核

Atomic Claim 18/574 · 2026-02-16_inferencex-v2-nvidia-blackwell-vs::IX2-0017

Claim: InferenceX 計畫在 2026 年稍晚加入 TPUv7 Ironwood。
Frame: ATTRIBUTE · Mode: ASSERTED · Mapping: PARTIAL
開啟逐條審核

Atomic Claim 19/574 · 2026-02-16_inferencex-v2-nvidia-blackwell-vs::IX2-0018

Claim: InferenceX 計畫在 2026 年稍晚加入 Trainium3
Frame: ATTRIBUTE · Mode: ASSERTED · Mapping: PARTIAL
開啟逐條審核

image

Source: InferenceMAX GitHub

Key Observations and Results to Highlight

We see competitive perf per TCO results on FP8 MI355X disagg+wideEP SGLang on AMD compared to FP8 B200 disagg+wideEP SGLang, but when compared to widely used Dynamo TRTLLM B200 FP8, TRT continues to framemog. This is amazing news that AMD SGLang Disagg prefill+wideEP for FP8 is able to match NVIDIA’s SGLang performance.

AMD 的 FP8 MI355X disagg+wideEP SGLang,相較 NVIDIA FP8 B200 disagg+wideEP SGLang,perf per TCO 已具有競爭力;但若拿來和大量實際部署的 Dynamo TRTLLM B200 FP8 比,TRT 仍然繼續 framemog、明顯領先。不過 AMD SGLang 的 FP8 Disagg prefill+wideEP 已經能追平 NVIDIA SGLang performance,這其實是非常好的消息。

Atomic Claim 20/574 · 2026-02-16_inferencex-v2-nvidia-blackwell-vs::IX2-0019

Claim:AMD 上以 SGLang 執行 FP8 MI355X disagg+wideEP,其每單位 TCO 效能相較 B200 FP8 disagg+wideEP SGLang 已具競爭力。
Frame: COMPARISON · Mode: ASSERTED · Mapping: COMPLETE
開啟逐條審核

Atomic Claim 21/574 · 2026-02-16_inferencex-v2-nvidia-blackwell-vs::IX2-0020

Claim: 但若拿來與廣泛使用的 Dynamo TRTLLM B200 FP8 比較,TRT 仍明顯領先。
Frame: NARY_RELATION · Mode: ASSERTED · Mapping: COMPLETE
開啟逐條審核

Atomic Claim 22/574 · 2026-02-16_inferencex-v2-nvidia-blackwell-vs::IX2-0021

Claim: 值得注意的是,AMDSGLang 執行 FP8 的 Disagg prefill+wideEP,效能已能匹配 NVIDIASGLang
Frame: ATTRIBUTE · Mode: ASSERTED · Mapping: PARTIAL
開啟逐條審核

We also see that for single node aggregated serving, AMD’s SGLang delivers better perf per TCO than NVIDIA’s SGLang for FP8. It is also great to see that AMD has deprecated their second class fork of vllm to move further upstream and closer to delivering first class experience. Stay tuned for our “State of AMD” article where we talk about the many areas where AMD’s pace of improvement has been rapid & also the areas where the pace of improvement has been lackluster. We recommend that NVIDIA focus even more on SGLang & vLLM ecosystem in addition their TRTLLM engine. Jensen needs to staff more resources & engineers towards contributing open ecosystems like SGLang & vLLM .

在 single-node aggregated serving 上,我們也看到 AMD SGLang 的 FP8 perf per TCO 優於 NVIDIA SGLang。另一個好消息是,AMD 已經淘汰那套 second-class 的 vLLM fork,往 upstream 靠攏,離 first-class user experience 更近一步 ↗。之後的《State of AMD》文章會談哪些地方 AMD 進步非常快、哪些地方仍然偏慢。我們也建議 NVIDIA 除了 TRTLLM engine 之外,更積極投入 SGLang 與 vLLM ecosystem;Jensen 應該增加更多資源與工程師,去貢獻 SGLang、vLLM 這類 open ecosystem ↗。

Atomic Claim 23/574 · 2026-02-16_inferencex-v2-nvidia-blackwell-vs::IX2-0022

Claim: 在 single-node aggregated serving 中,AMDSGLangFP8 下,每單位 TCO 效能優於 NVIDIASGLang
Frame: COMPARISON · Mode: ASSERTED · Mapping: COMPLETE
開啟逐條審核

Atomic Claim 24/574 · 2026-02-16_inferencex-v2-nvidia-blackwell-vs::IX2-0023

Claim: AMD 已淘汰自家較次級的 vLLM fork,轉向更接近 upstream,以改善使用體驗。
Frame: ATTRIBUTE · Mode: ASSERTED · Mapping: PARTIAL
開啟逐條審核

Atomic Claim 25/574 · 2026-02-16_inferencex-v2-nvidia-blackwell-vs::IX2-0024

Claim: SemiAnalysis 建議 NVIDIA 除了 TRTLLM engine 外,也應投入更多資源到 SGLangvLLM ecosystem。
Frame: ATTRIBUTE · Mode: ASSERTED · Mapping: COMPLETE
開啟逐條審核

Atomic Claim 26/574 · 2026-02-16_inferencex-v2-nvidia-blackwell-vs::IX2-0025

Claim: SemiAnalysis 認為 Jensen 應配置更多資源與工程師,投入 SGLangvLLM 等開放 ecosystem。
Frame: RELATION · Mode: ASSERTED · Mapping: COMPLETE
開啟逐條審核

SemiAnalysis InferenceX is free open source software and reader-supported. To receive new posts and support our work consider becoming a free or paid subscriber.

SemiAnalysis InferenceX 是免費 open-source software,也由讀者支持。如果想收到新文章並支持我們,可以考慮成為免費或付費訂閱者。

Subscribed

When it comes to the latest inference techniques that are used by the most prominent frontier large-scale inference services (such as disagg prefill+wideEP+FP4), Nvidia absolutely frame mogs with the B200, B300 and ASU frat leader, rack scale GB200/GB300 NVL72 across both SGLang and TRTLLM. Nvidia GPUs also dominate when it comes to energy efficiency, with much lower all-in provisioned picoJoules of energy per token across all workloads.

若看最先進、已被大型 frontier inference service 實際採用的技術組合,例如 disagg prefill+wideEP+FP4,Nvidia 在 SGLang 與 TRTLLM 兩邊都直接 framemog:B200、B300,以及堪稱 ASU frat leader 的 rack-scale GB200/GB300 NVL72 都非常強。Nvidia GPU 在 energy efficiency 上也占明顯優勢,所有 workload 的 all-in provisioned picoJoules per token 都更低。

Atomic Claim 27/574 · 2026-02-16_inferencex-v2-nvidia-blackwell-vs::IX2-0026

Claim: 在 frontier 大規模推論服務採用的最新技術(例如 disagg prefill+wideEP+FP4)上,NvidiaB200B300、rack-scale GB200GB300 NVL72,無論搭配 SGLangTRTLLM 都明顯領先。
Frame: NARY_RELATION · Mode: ASSERTED · Mapping: PARTIAL
開啟逐條審核

Atomic Claim 28/574 · 2026-02-16_inferencex-v2-nvidia-blackwell-vs::IX2-0027

Claim: Nvidia GPUs 在能源效率上也居領先,各種 workload 的 all-in provisioned picoJoules per token 都明顯較低。
Frame: COMPARISON · Mode: ASSERTED · Mapping: COMPLETE
開啟逐條審核

Turning to AMD, we find that the biggest issue with inference on their systems and using their software is composability . That is, many of AMDs inference optimization implementations work well in isolation, but when combined with other optimizations, the result is not as competitive as one would expect. Specifically, the composability of disagg prefill, wideEP and FP4 inference optimizations needs significant improvement.

轉到 AMD,我們認為目前 inference system 與 software 最大的問題是 composability ↗。也就是說,AMD 很多 inference optimization 單獨拿出來時表現不錯,但一旦和其他 optimization 疊在一起,結果就沒有想像中競爭。尤其 disagg prefill、wideEP 與 FP4 inference optimization 三者的 composability,還需要大幅改善。

Atomic Claim 29/574 · 2026-02-16_inferencex-v2-nvidia-blackwell-vs::IX2-0028

Claim: AMD 的許多推論最佳化,在單獨啟用時表現良好。
Frame: ATTRIBUTE · Mode: ASSERTED · Mapping: PARTIAL
開啟逐條審核

Atomic Claim 30/574 · 2026-02-16_inferencex-v2-nvidia-blackwell-vs::IX2-0029

Claim: 關於 AMD:但當多種最佳化組合在一起時,實際結果不如預期具有競爭力。
Frame: ATTRIBUTE · Mode: ASSERTED · Mapping: COMPLETE
開啟逐條審核

Atomic Claim 31/574 · 2026-02-16_inferencex-v2-nvidia-blackwell-vs::IX2-0030

Claim: 尤其 disagg prefill、wideEP 與 FP4 inference 等最佳化之間的 composability 仍需要大幅改善。
Frame: ATTRIBUTE · Mode: ASSERTED · Mapping: COMPLETE
開啟逐條審核

While performance is competitive on AMD when enabling just a subset of the SOTA inference optimizations, enabling all three major optimizations that labs use, AMD’s performance is currently not competitive with Nvidia’s. We strongly recommend to AMD that they focus heavily on composability of different inference optimizations. We have been told that AMD will start focusing on software composability of FP4+distributed inferencing across their whole software stack. This will happen after Chinese New Year as most of their disagg prefill+wideEP 10x inference engineers are based in China

AMD 若只開啟部分 SOTA inference optimization,performance 已經相當有競爭力;但把目前 lab 實際使用的三大 optimization 全部一起打開,整體 performance 還無法和 Nvidia 競爭。我們強烈建議 AMD 把不同 inference optimization 的 composability 當成核心優先事項。我們得知 AMD 在農曆年後會開始全面處理 FP4 + distributed inference 的 software composability,因為多數負責 disagg prefill+wideEP、能帶來 10x inference 改善的工程師都在中國。

Atomic Claim 32/574 · 2026-02-16_inferencex-v2-nvidia-blackwell-vs::IX2-0031

Claim: AMD 在只啟用部分 SOTA 推論最佳化時效能具競爭力,但當同時啟用 AI labs 常用的三項主要最佳化後,AMD 目前的效能仍無法與 Nvidia 競爭。
Frame: ATTRIBUTE · Mode: ASSERTED · Mapping: COMPLETE
開啟逐條審核

Atomic Claim 33/574 · 2026-02-16_inferencex-v2-nvidia-blackwell-vs::IX2-0032

Claim: SemiAnalysis 強烈建議 AMD 應高度聚焦不同推論最佳化之間的 composability。
Frame: ATTRIBUTE · Mode: ASSERTED · Mapping: COMPLETE
開啟逐條審核

Atomic Claim 34/574 · 2026-02-16_inferencex-v2-nvidia-blackwell-vs::IX2-0033

Claim: SemiAnalysis 得知,AMD 將開始強化整個 software stackFP4 + distributed inferencing 的 software composability。
Frame: ATTRIBUTE · Mode: RUMORED · Mapping: PARTIAL
開啟逐條審核

Atomic Claim 35/574 · 2026-02-16_inferencex-v2-nvidia-blackwell-vs::IX2-0034

Claim: 這項工作預計在農曆新年後展開,因為 AMD 多數負責 disagg prefill+wideEP 的核心 inference engineers 位於中國。
Frame: ATTRIBUTE · Mode: ASSERTED · Mapping: PARTIAL
開啟逐條審核

Nvidia’s GB300 NVL72 doesn’t disappoint. It achieves up to 100x on FP8 vs FP4 compared to even a strong H100 disagg+wideEP+MTP baseline and 65x on FP8 vs FP8. On H100 vs GB200 NVL72, we see up to 55x realized performance difference at 75 tok/s/user. Rack scale Blackwell NVL72 is framemogging hopper and makes hopper looks like it is jestermaxxing. As Jensen said at GTC 2025, he is chief revenue destroyer.

Nvidia GB300 NVL72 沒有令人失望。即使拿相當強的 H100 disagg+wideEP+MTP 當 baseline,GB300 在 FP8 對 FP4 的比較可達最高 100x、FP8 對 FP8 也可達 65x。H100 對 GB200 NVL72,在 75 tok/s/user 下,我們看到最高約 55x 的 realized performance 差距。Rack-scale Blackwell NVL72 正在 framemog Hopper,把 Hopper 襯得像在 jestermaxxing。正如 Jensen 在 GTC 2025 說的,他是 chief revenue destroyer ↗。

Atomic Claim 36/574 · 2026-02-16_inferencex-v2-nvidia-blackwell-vs::IX2-0035

Claim: NvidiaGB300 NVL72 效能表現符合甚至超越期待。
Frame: ATTRIBUTE · Mode: ASSERTED · Mapping: PARTIAL
開啟逐條審核

Atomic Claim 37/574 · 2026-02-16_inferencex-v2-nvidia-blackwell-vs::IX2-0036

Claim: 相較強勁的 H100 disagg+wideEP+MTP baseline,在 FP8FP4 的比較中最高可達 100 倍;另一個 FP8FP8 的比較則可達 65 倍。
Frame: COMPARISON · Mode: ASSERTED · Mapping: COMPLETE
開啟逐條審核

Atomic Claim 38/574 · 2026-02-16_inferencex-v2-nvidia-blackwell-vs::IX2-0037

Claim: 比較 H100GB200 NVL72,在 75 tok/s/user 下可觀察到最高 55 倍的實際效能差距。
Frame: COMPARISON · Mode: ASSERTED · Mapping: COMPLETE
開啟逐條審核

Atomic Claim 39/574 · 2026-02-16_inferencex-v2-nvidia-blackwell-vs::IX2-0038

Claim: Rack-scale Blackwell NVL72 的效能遠超 hopper,使 hopper 相形失色。
Frame: ATTRIBUTE · Mode: ASSERTED · Mapping: PARTIAL
開啟逐條審核

Atomic Claim 40/574 · 2026-02-16_inferencex-v2-nvidia-blackwell-vs::IX2-0039

Claim: 如 Jensen 在 GTC 2025 所說,他把自己稱為「chief revenue destroyer」。
Frame: CLAIM_ONLY · Mode: ATTRIBUTED · Mapping: CLAIM_ONLY
開啟逐條審核

At GTC 2024, Jensen claimed that Blackwell will deliver up to 30x perf on inference compared to H100, Jensen under promised & overdelivered on Blackwell inference performance. This should curtail the instances of analysts cracking “Jensen Math” jokes for some time.

GTC 2024 時 Jensen 宣稱 Blackwell inference performance 相較 H100 最多可提升 30x;最後結果是他在 Blackwell inference 上反而 under-promise、over-deliver。分析師拿「Jensen Math」開玩笑的次數,短期內應該可以少一點了。

Atomic Claim 41/574 · 2026-02-16_inferencex-v2-nvidia-blackwell-vs::IX2-0040

Claim: Jensen 在 GTC 2024 宣稱,Blackwell 的推論效能最高可達 H100 的 30 倍。
Frame: COMPARISON · Mode: ATTRIBUTED · Mapping: COMPLETE
開啟逐條審核

Atomic Claim 42/574 · 2026-02-16_inferencex-v2-nvidia-blackwell-vs::IX2-0040::SPLIT02

Claim: SemiAnalysis 認為,Jensen 對 Blackwell 推論效能是 underpromise、overdeliver。
Frame: ATTRIBUTE · Mode: INFERRED · Mapping: PARTIAL
開啟逐條審核

Atomic Claim 43/574 · 2026-02-16_inferencex-v2-nvidia-blackwell-vs::IX2-0041

Claim: 這應該會讓分析師一段時間內少開一些「Jensen Math」的玩笑。
Frame: CLAIM_ONLY · Mode: EXPECTED · Mapping: CLAIM_ONLY
開啟逐條審核

image

Source: SemiAnalysis InferenceX

Acknowledgments and InferenceX™ (formerly InferenceMAX) Initiative Supporters

We would like to thank Jensen Huang and Ian Buck for supporting this open-source effort by providing access to the latest GB300 NVL72 systems along with access to servers representing all GPU SKUs that they have produced for the past four years. We would like to thank the Nvidia team for allowing us to conduct independent benchmarks across this close to 1000 GPUs. Thank you to Jatin Gangani, Kedar Potdar, Sridhar Ramaswamy, Ishan Dhanani, Sahithi Chigurupati, along with many other Nvidia inference engineers for helping to validate and optimize Blackwell & Hopper configurations.

我們要感謝 Jensen Huang 與 Ian Buck 支持這項 open-source 工作,提供最新 GB300 NVL72 system,以及涵蓋 Nvidia 過去四年所有 GPU SKU 的 server access。也感謝 Nvidia team 讓我們能在接近 1,000 顆 GPU 上進行 independent benchmark;並謝謝 Jatin Gangani、Kedar Potdar、Sridhar Ramaswamy、Ishan Dhanani、Sahithi Chigurupati,以及許多 Nvidia inference engineer 協助驗證與最佳化 Blackwell、Hopper configuration。

We’re also grateful to Lisa Su and Anush Elangovan for their support of InferenceMAX and for supporting our work with the dozens of AMD engineers like Chun, Andy, Bill, Ramine, Theresa, Parth, etc that contributed to InferenceMAX & upstream vLLM/SGLang bug fixes, as well as for their responsiveness on helping debug and triage AMD exclusive bugs so as to help optimize AMD performance.

我們也感謝 Lisa Su 與 Anush Elangovan 對 InferenceMAX 的支持,以及協助我們和數十位 AMD engineer 合作,包括 Chun、Andy、Bill、Ramine、Theresa、Parth 等人。他們不只貢獻 InferenceMAX 與 upstream vLLM/SGLang bug fix,也在 AMD-exclusive bug 的 debug、triage 上快速回應,協助把 AMD performance 往上推。

We also want to recognize the SGLang, vLLM, and TensorRT-LLM maintainers for building a world-class software stack and open sourcing it to the entire world. You can check their articles on InferenceX here:

我們也要肯定 SGLang、vLLM、TensorRT-LLM 的 maintainer,打造世界級 software stack 並將其 open source 給全球使用。以下是他們在 InferenceX 上的相關文章:

SemiAnalysis InferenceMAX: vLLM maintainers & NVIDIA accelerate Blackwell Inference

SemiAnalysis InferenceMAX:vLLM maintainer 與 NVIDIA 共同加速 Blackwell Inference ↗

GPT-OSS Performance Optimizations: Pushing Pareto Frontier

GPT-OSS Performance Optimization:推進 Pareto Frontier ↗

SGLang & NVIDIA Accelerating SemiAnalysis InferenceMAX & GB200 Together

SGLang 與 NVIDIA 共同加速 SemiAnalysis InferenceMAX 與 GB200 ↗

The InferenceX initiative is also supported by many major buyers of compute and prominent members of the ML community including those from OpenAI, Microsoft, vLLM, Tri Dao, PyTorch Foundation, Oracle and more. You can find the full list here .

InferenceX initiative 也獲得許多大型 compute 買家與 ML community 重要成員支持,包括 OpenAI、Microsoft、vLLM、Tri Dao、PyTorch Foundation、Oracle 等。完整名單可在此查看 ↗。

SemiAnalysis InferenceX is free open source software and reader-supported. To receive new posts and support our work, consider becoming a free or paid subscriber.

SemiAnalysis InferenceX 是免費 open-source software,也由讀者支持。如果想收到新文章並支持我們,可以考慮成為免費或付費訂閱者。

Subscribed

A Primer on Important Technical Concepts

In this section, we will give a brief primer on technical concepts that may help the reader better interpret results. Some readers may not need this and can skip directly to our analysis of results. We will take a deeper dive into some of these topics after the results analysis.

這一節會先快速介紹幾個 technical concept,幫助讀者更容易解讀後面的 benchmark result。不需要這些基礎說明的讀者,可以直接跳到 results analysis;結果分析之後,我們還會再深入討論其中幾個主題。

Atomic Claim 44/574 · 2026-02-16_inferencex-v2-nvidia-blackwell-vs::IX2-0042

Claim: 本節會先簡要介紹一些有助於理解 benchmark 結果的技術概念。
Frame: CLAIM_ONLY · Mode: ASSERTED · Mapping: CLAIM_ONLY
開啟逐條審核

Atomic Claim 45/574 · 2026-02-16_inferencex-v2-nvidia-blackwell-vs::IX2-0044

Claim: 熟悉相關概念的讀者可以直接跳到結果分析。
Frame: CLAIM_ONLY · Mode: ASSERTED · Mapping: CLAIM_ONLY
開啟逐條審核

Atomic Claim 46/574 · 2026-02-16_inferencex-v2-nvidia-blackwell-vs::IX2-0043

Claim: 部分讀者可能不需要這段基礎介紹。
Frame: CLAIM_ONLY · Mode: HYPOTHETICAL · Mapping: CLAIM_ONLY
開啟逐條審核

Atomic Claim 47/574 · 2026-02-16_inferencex-v2-nvidia-blackwell-vs::IX2-0045

Claim: 在結果分析之後,文章還會對其中一些主題做更深入說明。
Frame: CLAIM_ONLY · Mode: ASSERTED · Mapping: CLAIM_ONLY
開啟逐條審核

Interactivity vs Throughput Tradeoff

The fundamental tradeoff with LLM inference is throughput versus latency. Interactivity (tok/s/user) describes how fast each user of a system receives tokens – it is the inverse of time per output token (TPOT). Throughput (tok/s) describes how many total tokens a system can crank out across all users. One can achieve higher total throughput by batching requests, but each request will be allocated less FLOPs and thus complete slower. This is analogous to the choice of riding a metro bus vs a race car. The metro bus serves many riders, but also makes frequent stops which takes time, but the cost of the metro bus can be amortized across many passengers. The race car can only carry one or two passengers, but it will make few if any additional stops meaning a faster travel time overall, but it is much more expensive to ride per passenger. The metro bus might make more sense for people heading to the park on a weekend, while the race car might be better for bringing a celebrity to their destination. There is no one size fits all solution.

LLM inference 最根本的 trade-off 是 throughput 與 latency。Interactivity(tok/s/user)描述單一使用者收到 token 的速度,也就是 time per output token(TPOT)的倒數;Throughput(tok/s)則是整套 system 對所有使用者合計每秒能產生多少 token。透過 batching 可以提高 total throughput,但每個 request 能分到的 FLOPs 變少,因此完成得更慢。這很像搭市區公車和賽車的差別:公車一次服務很多人,但一路停靠,花的時間較久,不過成本能分攤給大量乘客;賽車只能載一兩個人,幾乎不用停,因此總行程快很多,但每位乘客成本也高得多。週末去公園的人也許適合搭公車,送明星趕行程則可能適合賽車。不存在一套 configuration 可以適用所有情境。

Atomic Claim 48/574 · 2026-02-16_inferencex-v2-nvidia-blackwell-vs::IX2-0048

Claim: Throughput(tok/s)代表整個系統跨所有使用者每秒總共能產生多少 token。
Frame: CLAIM_ONLY · Mode: ASSERTED · Mapping: CLAIM_ONLY
開啟逐條審核

Atomic Claim 49/574 · 2026-02-16_inferencex-v2-nvidia-blackwell-vs::IX2-0046

Claim: LLM 推論最基本的 trade-off 是吞吐量與延遲之間的取捨。
Frame: ATTRIBUTE · Mode: ASSERTED · Mapping: COMPLETE
開啟逐條審核

Atomic Claim 50/574 · 2026-02-16_inferencex-v2-nvidia-blackwell-vs::IX2-0047

Claim: Interactivity(tok/s/user)代表系統中每位使用者收到 token 的速度,等同 time per output token(TPOT)的倒數。
Frame: ATTRIBUTE · Mode: ASSERTED · Mapping: COMPLETE
開啟逐條審核

Atomic Claim 51/574 · 2026-02-16_inferencex-v2-nvidia-blackwell-vs::IX2-0049

Claim: 透過 batching requests 可以提高整體吞吐量。
Frame: CLAIM_ONLY · Mode: ASSERTED · Mapping: CLAIM_ONLY
開啟逐條審核

Atomic Claim 52/574 · 2026-02-16_inferencex-v2-nvidia-blackwell-vs::IX2-0050

Claim: 但每個 request 分配到的 FLOPs 會變少,因此完成速度會更慢。
Frame: CLAIM_ONLY · Mode: ASSERTED · Mapping: CLAIM_ONLY
開啟逐條審核

Atomic Claim 53/574 · 2026-02-16_inferencex-v2-nvidia-blackwell-vs::IX2-0051

Claim: 這可類比成搭乘大眾巴士與賽車之間的選擇。
Frame: CLAIM_ONLY · Mode: ASSERTED · Mapping: CLAIM_ONLY
開啟逐條審核

Atomic Claim 54/574 · 2026-02-16_inferencex-v2-nvidia-blackwell-vs::IX2-0052

Claim: 大眾巴士一次可以服務許多乘客。
Frame: CLAIM_ONLY · Mode: ASSERTED · Mapping: CLAIM_ONLY
開啟逐條審核

Atomic Claim 55/574 · 2026-02-16_inferencex-v2-nvidia-blackwell-vs::IX2-0053

Claim: 但巴士會頻繁停靠,因此需要更多時間。
Frame: CLAIM_ONLY · Mode: ASSERTED · Mapping: CLAIM_ONLY
開啟逐條審核

Atomic Claim 56/574 · 2026-02-16_inferencex-v2-nvidia-blackwell-vs::IX2-0054

Claim: 巴士的固定成本可以由大量乘客共同攤提。
Frame: CLAIM_ONLY · Mode: ASSERTED · Mapping: CLAIM_ONLY
開啟逐條審核

Atomic Claim 57/574 · 2026-02-16_inferencex-v2-nvidia-blackwell-vs::IX2-0056

Claim: 賽車幾乎不用額外停靠,因此整體移動時間更短。
Frame: CLAIM_ONLY · Mode: ASSERTED · Mapping: CLAIM_ONLY
開啟逐條審核

Atomic Claim 58/574 · 2026-02-16_inferencex-v2-nvidia-blackwell-vs::IX2-0057

Claim: 但以每位乘客計算,賽車的成本明顯更高。
Frame: CLAIM_ONLY · Mode: ASSERTED · Mapping: CLAIM_ONLY
開啟逐條審核

Atomic Claim 59/574 · 2026-02-16_inferencex-v2-nvidia-blackwell-vs::IX2-0055

Claim: 賽車通常只能搭載一到兩名乘客。
Frame: CLAIM_ONLY · Mode: ASSERTED · Mapping: CLAIM_ONLY
開啟逐條審核

Atomic Claim 60/574 · 2026-02-16_inferencex-v2-nvidia-blackwell-vs::IX2-0058

Claim: 週末要去公園的大量乘客,可能更適合搭乘巴士。
Frame: CLAIM_ONLY · Mode: HYPOTHETICAL · Mapping: CLAIM_ONLY
開啟逐條審核

Atomic Claim 61/574 · 2026-02-16_inferencex-v2-nvidia-blackwell-vs::IX2-0059

Claim: 若是要快速把名人送到目的地,賽車可能更合適。
Frame: CLAIM_ONLY · Mode: HYPOTHETICAL · Mapping: CLAIM_ONLY
開啟逐條審核

Atomic Claim 62/574 · 2026-02-16_inferencex-v2-nvidia-blackwell-vs::IX2-0060

Claim: 不存在一種適合所有情境的單一最佳方案。
Frame: CLAIM_ONLY · Mode: ASSERTED · Mapping: CLAIM_ONLY
開啟逐條審核

image

Source: SemiAnalysis

Most benchmark results we will show in this article are InferenceX is a curve. It is important to analyze throughput at various levels of interactivity/latency instead of just looking at maximum achieved throughput (which normally can only be achieved at a single low interactivity). With inference, there is no one size fits all use case. The level of interactivity and throughput needed depends on the use case. For instance, real-time speech models require extremely low latency so that the end user can maintain a natural “conversation” with the LLM, whereas a basic QA chatbot may allow for higher latency. We leave it up to the reader to look at the curve and apply this principle to identify where their use case falls on the throughput-interactivity curve.

本文大多數 InferenceX benchmark result 都應該被視為一條 curve。重點不是只看 maximum throughput——因為最高 throughput 通常只能出現在很低 interactivity 的單一區域——而是要看不同 interactivity/latency 水準下能得到多少 throughput。Inference 沒有 one-size-fits-all use case;需要多少 interactivity 與 throughput,取決於應用本身。例如 real-time speech model 必須極低 latency,使用者才能和 LLM 保持自然對話;一般 QA chatbot 則可以容忍較高 latency。因此我們把判斷交給讀者,依自己的 use case 去看 throughput-interactivity curve 上真正相關的位置。

Atomic Claim 63/574 · 2026-02-16_inferencex-v2-nvidia-blackwell-vs::IX2-0061

Claim: 大多數 InferenceX benchmark 結果都以曲線呈現。
Frame: ATTRIBUTE · Mode: ASSERTED · Mapping: PARTIAL
開啟逐條審核

Atomic Claim 64/574 · 2026-02-16_inferencex-v2-nvidia-blackwell-vs::IX2-0062

Claim: 分析時應比較不同 interactivity/latency 水準下的 throughput,而不是只看最大吞吐量,因為最大吞吐量通常只能在單一低 interactivity 點達成。
Frame: CLAIM_ONLY · Mode: ASSERTED · Mapping: CLAIM_ONLY
開啟逐條審核

Atomic Claim 65/574 · 2026-02-16_inferencex-v2-nvidia-blackwell-vs::IX2-0063

Claim: 推論不存在適用所有 use case 的單一最佳配置。
Frame: CLAIM_ONLY · Mode: ASSERTED · Mapping: CLAIM_ONLY
開啟逐條審核

Atomic Claim 66/574 · 2026-02-16_inferencex-v2-nvidia-blackwell-vs::IX2-0064

Claim: 所需的 interactivity 與 throughput 取決於實際 use case。
Frame: CLAIM_ONLY · Mode: ASSERTED · Mapping: CLAIM_ONLY
開啟逐條審核

Atomic Claim 67/574 · 2026-02-16_inferencex-v2-nvidia-blackwell-vs::IX2-0065

Claim: 例如 real-time speech model 需要極低延遲,才能讓終端使用者與 LLM 維持自然的「對話」。
Frame: ATTRIBUTE · Mode: ASSERTED · Mapping: COMPLETE
開啟逐條審核

Atomic Claim 68/574 · 2026-02-16_inferencex-v2-nvidia-blackwell-vs::IX2-0067

Claim: 使用者應根據 benchmark curve 判斷自己的 use case 落在 throughput-interactivity curve 的哪個位置。
Frame: CLAIM_ONLY · Mode: ASSERTED · Mapping: CLAIM_ONLY
開啟逐條審核

The Cost/Perf per TCO vs Interactivity/End-to-End Latency curve mostly follows the Throughput vs Interactivity/End-to-End Latency Curve: More tokens/hour leads to a lower cost per token as fixed $/hour costs are amortized over more tokens produced.

Cost/Perf per TCO 對 Interactivity/End-to-End Latency 的 curve,大致會跟 Throughput 對 Interactivity/End-to-End Latency 的 curve 同方向:每小時產生的 token 越多,fixed $/hour cost 能分攤到更多 token 上,因此 cost per token 越低。

Atomic Claim 69/574 · 2026-02-16_inferencex-v2-nvidia-blackwell-vs::IX2-0068

Claim: Cost/Perf per TCO vs Interactivity/End-to-End Latency 曲線大致跟隨 Throughput vs Interactivity/End-to-End Latency 曲線:每小時產生的 token 越多,固定的每小時成本就能攤在更多 token 上,因此每 token 成本更低。
Frame: CLAIM_ONLY · Mode: ASSERTED · Mapping: CLAIM_ONLY
開啟逐條審核

Prefill and Decode Phases

Inference contains two main phases: prefill and decode. Prefill occurs during the first forward pass of a request’s lifetime. It is computationally intensive since all tokens in the request are processed in parallel. This phase is responsible for “filling up” the KV cache for a sequence. After prefill, responses are generated (or decoded) one token at a time. Each forward pass loads the entire KV cache for a sequence from HBM, while only performing the computation for a single token, making decode memory (bandwidth) intensive.

Inference 有兩個主要 phase:prefill 與 decode。Prefill 發生在一個 request 生命週期的第一次 forward pass;因為 request 內所有 token 會平行處理,所以 compute intensity 很高,這一階段同時負責把該 sequence 的 KV cache「填起來」。Prefill 完成後,response 進入 decode,一次產生一個 token。每次 forward pass 都得從 HBM 載入整個 sequence 的 KV cache,但只計算單一 token,因此 decode 本質上更吃 memory bandwidth。

Atomic Claim 70/574 · 2026-02-16_inferencex-v2-nvidia-blackwell-vs::IX2-0069

Claim: 推論流程包含 prefill
Frame: ATTRIBUTE · Mode: ASSERTED · Mapping: PARTIAL
開啟逐條審核

Atomic Claim 71/574 · 2026-02-16_inferencex-v2-nvidia-blackwell-vs::IX2-0070

Claim: 推論流程也包含 decode
Frame: ATTRIBUTE · Mode: ASSERTED · Mapping: PARTIAL
開啟逐條審核

Atomic Claim 72/574 · 2026-02-16_inferencex-v2-nvidia-blackwell-vs::IX2-0074

Claim: 完成 prefill 後,response 會逐 token 生成,也就是進入 decode
Frame: ATTRIBUTE · Mode: ASSERTED · Mapping: PARTIAL
開啟逐條審核

Atomic Claim 73/574 · 2026-02-16_inferencex-v2-nvidia-blackwell-vs::IX2-0071

Claim: Prefill 發生在一個 request 生命週期的第一次 forward pass。
Frame: ATTRIBUTE · Mode: ASSERTED · Mapping: COMPLETE
開啟逐條審核

Atomic Claim 74/574 · 2026-02-16_inferencex-v2-nvidia-blackwell-vs::IX2-0072

Claim: 由於 request 中所有 token 會平行處理,因此 prefill 屬於運算密集型階段。
Frame: ATTRIBUTE · Mode: ASSERTED · Mapping: PARTIAL
開啟逐條審核

Atomic Claim 75/574 · 2026-02-16_inferencex-v2-nvidia-blackwell-vs::IX2-0073

Claim: 此階段負責為一段 sequence 建立或「填滿」KV cache
Frame: ATTRIBUTE · Mode: ASSERTED · Mapping: PARTIAL
開啟逐條審核

Atomic Claim 76/574 · 2026-02-16_inferencex-v2-nvidia-blackwell-vs::IX2-0075

Claim: 每次 forward pass 都必須從 HBM 載入整段 sequence 的 KV cache,但只計算一個 token,因此 decode 屬於記憶體頻寬密集型工作。
Frame: ATTRIBUTE · Mode: ASSERTED · Mapping: COMPLETE
開啟逐條審核

When prefill and decode performed on the same engine, prefill constantly disrupts decode batches leading to worse overall performance.

當 prefill 與 decode 跑在同一組 engine 上時,prefill 會不斷打斷 decode batch,導致整體 performance 變差。

Atomic Claim 77/574 · 2026-02-16_inferencex-v2-nvidia-blackwell-vs::IX2-0076

Claim:prefilldecode 在同一個 engine 上執行時,prefill 會持續打斷 decode batches,造成整體效能變差。
Frame: ATTRIBUTE · Mode: ASSERTED · Mapping: COMPLETE
開啟逐條審核

Disaggregated Prefill

Disaggregated prefill (aka PD disaggregation or simply “disagg”) is the practice of separating the prefill and decode phases across separate pools of GPUs or clusters. These separate prefill and decode pools can be tuned independently and scaled to match the needs of workloads.

Disaggregated prefill(也稱 PD disaggregation,或簡稱 disagg)就是把 prefill 與 decode 分離到不同 GPU pool 或 cluster。這兩組 prefill/decode pool 可以分別 tuning,並依 workload 需求獨立 scale。

Atomic Claim 78/574 · 2026-02-16_inferencex-v2-nvidia-blackwell-vs::IX2-0077

Claim: Disaggregated prefill(也稱 PD disaggregation 或簡稱 disagg)是把 prefilldecode 階段分離到不同的 GPUs pools 或 clusters 的做法。
Frame: ATTRIBUTE · Mode: ASSERTED · Mapping: PARTIAL
開啟逐條審核

Atomic Claim 79/574 · 2026-02-16_inferencex-v2-nvidia-blackwell-vs::IX2-0078

Claim: 分離後的 prefilldecode pools 可以各自獨立調校,並依 workload 需求分別擴展。
Frame: ATTRIBUTE · Mode: ASSERTED · Mapping: PARTIAL
開啟逐條審核

Tensor Parallel, Expert Parallel, Data Parallel (TP, EP, DP)

TP allows for maximize interactivity at small batch sizes, but it must carry out an all-reduce at every layer. EP shards experts, exploiting MoE sparsity, with the drawback being an all-to-all collective (which is more costly than simpler collectives like all-reduce) is carried out for MoE layers and can be imbalanced at small batches. DP replicates the entire model (or just parts of a model, like attention) on multiple groups of GPUs (ranks) and then load balances requests among ranks. It is the simplest to scale, but repeats weight loading which can be wasteful at scale.

TP 在 small batch 下能最大化 interactivity,但每一層都必須做一次 all-reduce。EP 會把 expert shard 開來、利用 MoE sparsity,但代價是每個 MoE layer 都需要 all-to-all collective;all-to-all 比 all-reduce 等簡單 collective 更昂貴,而且 small batch 下還可能出現 load imbalance。DP 則把完整 model——或 attention 等部分 model——複製到多組 GPU rank,再把 request load balance 到不同 rank;它最容易 scale,但會重複載入 weight,大規模時可能浪費 bandwidth。

Atomic Claim 80/574 · 2026-02-16_inferencex-v2-nvidia-blackwell-vs::IX2-0080

Claim: 但 TP 每一層都必須執行一次 all-reduce
Frame: ATTRIBUTE · Mode: ASSERTED · Mapping: PARTIAL
開啟逐條審核

Atomic Claim 81/574 · 2026-02-16_inferencex-v2-nvidia-blackwell-vs::IX2-0079

Claim: TP 可在小 batch size 下最大化 interactivity。
Frame: ATTRIBUTE · Mode: ASSERTED · Mapping: COMPLETE
開啟逐條審核

Atomic Claim 82/574 · 2026-02-16_inferencex-v2-nvidia-blackwell-vs::IX2-0081

Claim: EP 會把 experts 分片,以利用 MoEsparsity
Frame: RELATION · Mode: ASSERTED · Mapping: COMPLETE
開啟逐條審核

Atomic Claim 83/574 · 2026-02-16_inferencex-v2-nvidia-blackwell-vs::IX2-0082

Claim: 其缺點是 MoE layers 需要執行 all-to-all collective,而這比 all-reduce 等較簡單 collective 成本更高。
Frame: COMPARISON · Mode: ASSERTED · Mapping: COMPLETE
開啟逐條審核

Atomic Claim 84/574 · 2026-02-16_inferencex-v2-nvidia-blackwell-vs::IX2-0084

Claim: DP 會在多組 GPUs(ranks)上複製整個模型,或只複製模型部分(例如 attention),再把 requests 在各 ranks 之間做 load balancing。
Frame: ATTRIBUTE · Mode: ASSERTED · Mapping: COMPLETE
開啟逐條審核

Tracking Improvements Over Time

One of the main goals of InferenceX is to visualize performance improvements over time. While new chips are released on an O(yearly) cadence, software releases happen on an O(weekly) cadence. Our goal is to constantly update recipes with the latest and greatest software improvements and benchmark the configurations.

InferenceX 的主要目標之一,就是把 performance 隨時間的改善視覺化。新 chip 大約是 O(yearly) 的更新頻率,software release 卻是 O(weekly)。我們希望持續把最新、最好的 software improvement 更新到 recipe,再重新 benchmark 各種 configuration。

Atomic Claim 85/574 · 2026-02-16_inferencex-v2-nvidia-blackwell-vs::IX2-0087

Claim: InferenceX 的主要目標之一,是視覺化效能隨時間的改善。
Frame: ATTRIBUTE · Mode: ASSERTED · Mapping: PARTIAL
開啟逐條審核

Atomic Claim 86/574 · 2026-02-16_inferencex-v2-nvidia-blackwell-vs::IX2-0088

Claim: 關於 InferenceX:新晶片大約以年度 cadence 推出,但 software releases 大約是每週 cadence。
Frame: ATTRIBUTE · Mode: ASSERTED · Mapping: COMPLETE
開啟逐條審核

Atomic Claim 87/574 · 2026-02-16_inferencex-v2-nvidia-blackwell-vs::IX2-0089

Claim: InferenceX 的目標是持續用最新 software improvements 更新 recipes,並重新 benchmark 各種 configurations。
Frame: ATTRIBUTE · Mode: ASSERTED · Mapping: COMPLETE
開啟逐條審核

DeepSeek R1

The AMD team has significantly improved performance for all configurations of SGLang DeepSeek R1 FP4. For the same interactivity, AMD has almost doubled the amount of throughput in the span of less than 2 months. Moreover, we have pushed AMD to upstream performance enhancing changes from their forked SGLang images into the official SGLang image. From December 2025 to January 2026, AMD’s software was improved up to 2x in performance.

AMD team 已大幅改善所有 SGLang DeepSeek R1 FP4 configuration 的 performance。在相同 interactivity 下,不到兩個月時間,AMD 的 throughput 幾乎翻倍。此外,我們也持續推 AMD 把 forked SGLang image 裡提升 performance 的改動 upstream 到 official SGLang image。從 2025 年 12 月到 2026 年 1 月,AMD software 的 performance 最高改善約 2x。

Atomic Claim 88/574 · 2026-02-16_inferencex-v2-nvidia-blackwell-vs::IX2-0090

Claim: AMD 團隊已顯著改善所有 SGLang DeepSeek R1 FP4 configurations 的效能。
Frame: ATTRIBUTE · Mode: ASSERTED · Mapping: PARTIAL
開啟逐條審核

Atomic Claim 89/574 · 2026-02-16_inferencex-v2-nvidia-blackwell-vs::IX2-0091

Claim: 在相同 interactivity 下,AMD 不到兩個月就幾乎把 throughput 提高一倍。
Frame: COMPARISON · Mode: ASSERTED · Mapping: PARTIAL
開啟逐條審核

Atomic Claim 90/574 · 2026-02-16_inferencex-v2-nvidia-blackwell-vs::IX2-0092

Claim: SemiAnalysis 也推動 AMD 把 forked SGLang images 中的效能改善 upstream 回官方 SGLang image。
Frame: ATTRIBUTE · Mode: ASSERTED · Mapping: PARTIAL
開啟逐條審核

Atomic Claim 91/574 · 2026-02-16_inferencex-v2-nvidia-blackwell-vs::IX2-0093

Claim: 從 2025 年 12 月到 2026 年 1 月,AMD software 效能最高改善接近 2 倍。
Frame: COMPARISON · Mode: ASSERTED · Mapping: PARTIAL
開啟逐條審核

image

Source: SemiAnalysis InferenceX

In order to continue becoming closer to an first class experience, AMD needs increase their support of vLLM & SGLang maintainers through compute contributions and code contributions & having more reviewers that work for AMD to speed up the review process of AMD PRs into the upstream.

若要繼續往 first-class experience 靠攏,AMD 還需要增加對 vLLM 與 SGLang maintainer 的支援,包括提供更多 compute contribution、code contribution,以及安排更多 AMD reviewer,加快 AMD PR 進 upstream 的 review 流程。

Atomic Claim 92/574 · 2026-02-16_inferencex-v2-nvidia-blackwell-vs::IX2-0094

Claim: 若要更接近 first-class experience,AMD 需要透過算力與程式碼貢獻,增加對 vLLMSGLang maintainers 的支援,並增加更多 AMD 內部 reviewers,以加快 AMD PRs 進 upstream 的 review 流程。
Frame: NARY_RELATION · Mode: HYPOTHETICAL · Mapping: PARTIAL
開啟逐條審核

image

Source: SemiAnalysis

On the other hand, Nvidia’s results were more consistent, with minor improvements for B200 SGLang over a similar time period.

相較之下,Nvidia 的 result 更穩定;同一段期間內,B200 SGLang 只有小幅改善。

Atomic Claim 93/574 · 2026-02-16_inferencex-v2-nvidia-blackwell-vs::IX2-0095

Claim: 另一方面,Nvidia 的結果較穩定,同期 B200 SGLang 只有小幅改善。
Frame: COMPARISON · Mode: ASSERTED · Mapping: PARTIAL
開啟逐條審核

image

Source: SemiAnalysis InferenceX

Many of the mature SKUs had minimal improvements. For example, H200 TRT single node has not changed in performance in the span of 4 months since October, but this is because Hopper support has been excellent since day 1, and performance has close to peak theoretical for this workload all along, making it hard to deliver incremental performance gains.

不少成熟 SKU 幾乎沒有再進步。例如 H200 TRT single-node 從 10 月到現在四個月 performance 都沒什麼變,但原因是 Hopper 從 Day 1 起 support 就已非常完整,而且這個 workload 的表現一直很接近 theoretical peak,因此後續本來就很難再擠出 incremental gain。

Atomic Claim 94/574 · 2026-02-16_inferencex-v2-nvidia-blackwell-vs::IX2-0096

Claim: 關於 DeepSeek R1:許多已成熟的 SKUs 幾乎沒有明顯改善。
Frame: ATTRIBUTE · Mode: ASSERTED · Mapping: COMPLETE
開啟逐條審核

Atomic Claim 95/574 · 2026-02-16_inferencex-v2-nvidia-blackwell-vs::IX2-0097

Claim: 例如 H200 TRT single-node 自 10 月起四個月內效能都沒有變化。
Frame: ATTRIBUTE · Mode: ASSERTED · Mapping: COMPLETE
開啟逐條審核

Atomic Claim 96/574 · 2026-02-16_inferencex-v2-nvidia-blackwell-vs::IX2-0098

Claim: 原因是 Hopper 從第一天起 software support 就已非常成熟。
Frame: ATTRIBUTE · Mode: ASSERTED · Mapping: COMPLETE
開啟逐條審核

Atomic Claim 97/574 · 2026-02-16_inferencex-v2-nvidia-blackwell-vs::IX2-0099

Claim: 關於 H200:該 workload 的效能一直都已接近理論峰值,因此很難再取得 incremental performance gains。
Frame: ATTRIBUTE · Mode: ASSERTED · Mapping: PARTIAL
開啟逐條審核

image

Source: SemiAnalysis InferenceX

MI300X and MI325X have seen some improvements, mainly from the most recent SGLang release. Note that for much of the history of InferenceX, AMD was using “private” ROCm images that were not upstreamed, so runs prior to ~Jan 2026 cannot be compared directly to those that are more recent.

MI300X、MI325X 則有一些改善,主要來自最新一版 SGLang。要注意的是,在 InferenceX 很長一段歷史裡,AMD 使用的是沒有 upstream 的「private」ROCm image,因此大約 2026 年 1 月以前的 run,不能直接和最近 result 一比一比較。

Atomic Claim 98/574 · 2026-02-16_inferencex-v2-nvidia-blackwell-vs::IX2-0100

Claim: MI300XMI325X 也有一些改善,主要來自最新的 SGLang release。
Frame: ATTRIBUTE · Mode: ASSERTED · Mapping: PARTIAL
開啟逐條審核

Atomic Claim 99/574 · 2026-02-16_inferencex-v2-nvidia-blackwell-vs::IX2-0101

Claim: 需注意,在 InferenceX 很長一段歷史期間,AMD 使用的是尚未 upstream 的「private」ROCm images,因此約 2026 年 1 月以前的測試結果不能直接與近期結果比較。
Frame: NARY_RELATION · Mode: ASSERTED · Mapping: PARTIAL
開啟逐條審核

image

Source: SemiAnalysis InferenceX

GB200 Dynamo TRT-LLM disagg has seen some significant improvements as well, with a 20% increase in max throughput in the span of a little over 1 month. We also see improvements in the middle interactivities, where wide EP is deployed. This is likely due to maturing wide EP kernels on GB200.

GB200 Dynamo TRT-LLM disagg 也有明顯進步,一個多月內 max throughput 增加約 20%。在會部署 wide EP 的中間 interactivity 區域,我們同樣看到改善;這很可能來自 GB200 wide EP kernel 持續成熟。

Atomic Claim 100/574 · 2026-02-16_inferencex-v2-nvidia-blackwell-vs::IX2-0102

Claim: GB200 Dynamo TRT-LLM disagg 也有顯著改善。
Frame: ATTRIBUTE · Mode: ASSERTED · Mapping: PARTIAL
開啟逐條審核

Atomic Claim 101/574 · 2026-02-16_inferencex-v2-nvidia-blackwell-vs::IX2-0103

Claim: GB200 的最大 throughput 在略超過一個月內提高約 20%。
Frame: ATTRIBUTE · Mode: ASSERTED · Mapping: COMPLETE
開啟逐條審核

Atomic Claim 102/574 · 2026-02-16_inferencex-v2-nvidia-blackwell-vs::IX2-0104

Claim: 在採用 wide EP 的中等 interactivity 區間,也可以看到效能改善。
Frame: ATTRIBUTE · Mode: ASSERTED · Mapping: COMPLETE
開啟逐條審核

Atomic Claim 103/574 · 2026-02-16_inferencex-v2-nvidia-blackwell-vs::IX2-0105

Claim: 這很可能來自 GB200 wide EP kernels 日益成熟。
Frame: ATTRIBUTE · Mode: INFERRED · Mapping: PARTIAL
開啟逐條審核

image

Source: SemiAnalysis InferenceX

B200 SGLang has seen steady and continuous improvement for both FP4 and FP8 scenarios since our initial launch, with throughput per GPU doubling at some interactivity levels since last October.

自 InferenceX 初次發布以來,B200 SGLang 在 FP4、FP8 scenario 都持續穩定進步;和去年 10 月相比,某些 interactivity level 的 throughput per GPU 已經翻倍。

Atomic Claim 104/574 · 2026-02-16_inferencex-v2-nvidia-blackwell-vs::IX2-0109

Claim: 自去年 10 月以來,在部分 interactivity 水準,B200 SGLang 的每 GPU throughput 已提高一倍。
Frame: COMPARISON · Mode: ASSERTED · Mapping: COMPLETE
開啟逐條審核

Atomic Claim 105/574 · 2026-02-16_inferencex-v2-nvidia-blackwell-vs::IX2-0107

Claim: 自首次發布以來,B200 SGLangFP4 情境下持續改善。
Frame: NARY_RELATION · Mode: ASSERTED · Mapping: COMPLETE
開啟逐條審核

Atomic Claim 106/574 · 2026-02-16_inferencex-v2-nvidia-blackwell-vs::IX2-0106

Claim: B200 SGLang 的效能持續穩定改善。
Frame: ATTRIBUTE · Mode: ASSERTED · Mapping: COMPLETE
開啟逐條審核

Atomic Claim 107/574 · 2026-02-16_inferencex-v2-nvidia-blackwell-vs::IX2-0108

Claim: B200 SGLangFP8 情境下也持續改善。
Frame: ATTRIBUTE · Mode: ASSERTED · Mapping: PARTIAL
開啟逐條審核

image

Source: SemiAnalysis InferenceX

For MI355X Disaggregated inference serving, AMD recommends using SGLang with MoRI. MoRI is AMD’s MoE dispatch/combine collective and KV Cache transfer library built from first principles by AMD’s cracked 10x China-based engineering team. Although MoRI needs much more open CI and testing, we are strong supporters of the direction that MoRI is taking. This is because instead of taking AMD’s historical approach, which was to fork NVIDIA’s NCCL into RCCL, MoRI is built from scratch by taking the lessons from RCCL/NCCL and building an entirely new package from first principles. The use of MoRI has also delivered good speedups in the span of more than a month, with throughput per GPU increasing by more than 20% in the 20-45 tok/s/user interactivity range.

MI355X 做 Disaggregated inference serving 時,AMD 建議使用 SGLang + MoRI。MoRI 是 AMD 從 first principles 自行打造的 MoE dispatch/combine collective 與 KV Cache transfer library ↗,由一支很強的中國 10x engineering team 開發。MoRI 還需要更多 open CI 與 testing,但我們非常支持這個方向,因為它不像 AMD 過去把 NVIDIA NCCL fork 成 RCCL,而是吸收 RCCL/NCCL 的經驗後,重新從零打造一套新 package。過去一個多月,MoRI 也帶來不錯 speedup:在 20–45 tok/s/user interactivity 範圍,throughput per GPU 增加超過 20%。

Atomic Claim 108/574 · 2026-02-16_inferencex-v2-nvidia-blackwell-vs::IX2-0110

Claim:MI355X Disaggregated inference serving,AMD 建議搭配 MoRI 使用 SGLang
Frame: NARY_RELATION · Mode: ATTRIBUTED · Mapping: COMPLETE
開啟逐條審核

Atomic Claim 109/574 · 2026-02-16_inferencex-v2-nvidia-blackwell-vs::IX2-0111

Claim: MoRIAMD 從零打造的 MoE dispatch/combine collective 與 KV Cache transfer library,主要由 AMD 位於中國的核心工程團隊開發。
Frame: ATTRIBUTE · Mode: ASSERTED · Mapping: PARTIAL
開啟逐條審核

Atomic Claim 110/574 · 2026-02-16_inferencex-v2-nvidia-blackwell-vs::IX2-0112

Claim: 雖然 MoRI 還需要更多公開 CI 與測試,SemiAnalysis 對 MoRI 的發展方向高度支持。
Frame: ATTRIBUTE · Mode: ASSERTED · Mapping: COMPLETE
開啟逐條審核

Atomic Claim 111/574 · 2026-02-16_inferencex-v2-nvidia-blackwell-vs::IX2-0113

Claim: 原因是 MoRI 不再沿用 AMD 過去把 NVIDIA NCCL fork 成 RCCL 的方式,而是吸收 RCCLNCCL 經驗後,從 first principles 重新打造一套全新 package。
Frame: ATTRIBUTE · Mode: ASSERTED · Mapping: PARTIAL
開啟逐條審核

Atomic Claim 112/574 · 2026-02-16_inferencex-v2-nvidia-blackwell-vs::IX2-0114

Claim: 在超過一個月的期間內,MoRI 也帶來明顯加速;在 20~45 tok/s/user interactivity 區間,每 GPU throughput 提升超過 20%。
Frame: ATTRIBUTE · Mode: ASSERTED · Mapping: COMPLETE
開啟逐條審核

image

Source: SemiAnalysis InferenceX

GPT-OSS 120B

For MI300X and MI325X, we have seen marginal improvements across the board. Some AITER optimizations helped MI300X performance across all interactivities, and switching to the upstream vLLM ROCm image led to improvements.

MI300X、MI325X 整體則是小幅改善。一些 AITER optimization 讓 MI300X 在各個 interactivity 都有所提升,而切換到 upstream vLLM ROCm image 也帶來額外進步。

Atomic Claim 113/574 · 2026-02-16_inferencex-v2-nvidia-blackwell-vs::IX2-0115

Claim:MI300XMI325X,各種 interactivity 下都只有小幅改善。
Frame: ATTRIBUTE · Mode: ASSERTED · Mapping: COMPLETE
開啟逐條審核

Atomic Claim 114/574 · 2026-02-16_inferencex-v2-nvidia-blackwell-vs::IX2-0116

Claim: 部分 AITER 最佳化改善了 MI300X 在所有 interactivity 下的效能,而切換至 upstream vLLM ROCm image 後也有進一步改善。
Frame: ATTRIBUTE · Mode: ASSERTED · Mapping: COMPLETE
開啟逐條審核

SemiAnalysis InferenceX is free open source software and reader-supported. To receive new posts and support our work, consider becoming a free or paid subscriber.

SemiAnalysis InferenceX 是免費 open-source software,也由讀者支持。如果想收到新文章並支持我們,可以考慮成為免費或付費訂閱者。

Subscribed

image

Source: SemiAnalysis InferenceX

In the case of the MI325X, it appears that not all performance enhancements that were present in the downstream ROCm fork image (used during the October 5th, 2025 run) have made it into the official vLLM ROCm image.
Unfortunately, the MI355X literally still uses a fork of the vLLM 0.10.1 build rocm/7.0:rocm7.0_ubuntu_22.04_vllm_0.10.1_instinct_20250927_rc1). We would love to have seen it updated it by now, but unfortunately the current official image (0.15.1, at the time this article was written) is not yet optimized for the MI355X and runs into hard errors. We had also run into hard errors crashes on Mi355 for vLLM 0.14. Word on the street is that vLLM 0.16.0 will finally deliver all the changes needed for better MI355X performance.

Atomic Claim 115/574 · 2026-02-16_inferencex-v2-nvidia-blackwell-vs::IX2-0117

Claim:MI325X 而言,2025 年 10 月 5 日測試時 downstream ROCm fork image 中的部分效能最佳化,似乎尚未全部進入官方 vLLM ROCm image。
Frame: ATTRIBUTE · Mode: ASSERTED · Mapping: PARTIAL
開啟逐條審核

Atomic Claim 116/574 · 2026-02-16_inferencex-v2-nvidia-blackwell-vs::IX2-0118

Claim: MI355X 目前仍在使用 fork 自 vLLM 0.10.1 的 build(rocm/7.0:rocm7.0_ubuntu_22.04_vllm_0.10.1_instinct_20250927_rc1)。
Frame: RELATION · Mode: ASSERTED · Mapping: PARTIAL
開啟逐條審核

Atomic Claim 117/574 · 2026-02-16_inferencex-v2-nvidia-blackwell-vs::IX2-0119

Claim: 關於 GPT-OSS:SemiAnalysis 原本希望這個版本現在已經能更新。
Frame: ATTRIBUTE · Mode: ASSERTED · Mapping: COMPLETE
開啟逐條審核

Atomic Claim 118/574 · 2026-02-16_inferencex-v2-nvidia-blackwell-vs::IX2-0120

Claim: 但撰文時的官方 image 0.15.1 尚未針對 MI355X 最佳化,而且會遇到 hard errors。
Frame: ATTRIBUTE · Mode: ASSERTED · Mapping: PARTIAL
開啟逐條審核

Atomic Claim 119/574 · 2026-02-16_inferencex-v2-nvidia-blackwell-vs::IX2-0121

Claim:MI355X 上執行 vLLM 0.14 時,SemiAnalysis 也遇過 hard errors 與 crashes。
Frame: ATTRIBUTE · Mode: ASSERTED · Mapping: PARTIAL
開啟逐條審核

Atomic Claim 120/574 · 2026-02-16_inferencex-v2-nvidia-blackwell-vs::IX2-0122

Claim: 市場消息指出,vLLM 0.16.0 最終會納入提升 MI355X 效能所需的全部變更。
Frame: ATTRIBUTE · Mode: RUMORED · Mapping: PARTIAL
開啟逐條審核

image

Source: SemiAnalysis InferenceX

Turning back to Nvidia’s systems, both Hopper and Blackwell saw a steady performance increase between vLLM 0.11.2 and 0.13.0. Soon, we will update recipes for Nvidia GPUs to use the latest vLLM version and we expect even greater performance gains after making the switch. We also observed a performance bump in the latest 1.2.0 version of TRT-LLM.

回到 Nvidia system,Hopper 與 Blackwell 從 vLLM 0.11.2 升到 0.13.0 都出現穩定 performance improvement。我們很快會把 Nvidia GPU recipe 更新到最新 vLLM version,預期切換後還會再看到更大的 performance gain。最新 TRT-LLM 1.2.0 也有明顯提升。

Atomic Claim 121/574 · 2026-02-16_inferencex-v2-nvidia-blackwell-vs::IX2-0123

Claim: 回到 Nvidia 系統,HopperBlackwellvLLM 0.11.2 升至 0.13.0 之間,都有穩定的效能提升。
Frame: ATTRIBUTE · Mode: ASSERTED · Mapping: COMPLETE
開啟逐條審核

Atomic Claim 122/574 · 2026-02-16_inferencex-v2-nvidia-blackwell-vs::IX2-0124

Claim: InferenceX 很快會把 Nvidia GPUs 的 recipes 更新至最新 vLLM 版本,並預期完成這項 switch 後還會有更大的效能改善。
Frame: ATTRIBUTE · Mode: EXPECTED · Mapping: PARTIAL
開啟逐條審核

Atomic Claim 123/574 · 2026-02-16_inferencex-v2-nvidia-blackwell-vs::IX2-0125

Claim: 最新 1.2.0 版 TRT-LLM 也觀察到效能提升。
Frame: ATTRIBUTE · Mode: ASSERTED · Mapping: PARTIAL
開啟逐條審核

image

Source: SemiAnalysis InferenceX

image

Source: SemiAnalysis InferenceX

Disaggregated Inference Frameworks

NVIDIA uses Dynamo for its disaggregated inference setup. Dynamo is an inference framework designed for multi-node distributed inference, featuring techniques such as prefill-decode disaggregation, request routing, and KV cache offloading. It is inference-engine agnostic, allowing us to use SGLang and TRT LLM as backends in our benchmark. For AMD, we use SGLang with two different KV cache transfer frameworks: MoRI and Mooncake. MoRI is a high-performance communication interface focusing on RDMA and GPU integration, offering applications such as network collective operations and expert parallel kernels. Mooncake, which recently joined the PyTorch ecosystem , supports prefill-decode disaggregation and many fault tolerant multi-node features.

NVIDIA 的 disaggregated inference setup 使用 Dynamo。Dynamo ↗ 是為 multi-node distributed inference 設計的 inference framework,包含 prefill-decode disaggregation、request routing、KV cache offloading 等技術;它與 inference engine 解耦,因此我們 benchmark 可以分別用 SGLang、TRT-LLM 當 backend。AMD 這邊則使用 SGLang,搭配兩種 KV cache transfer framework:MoRI 與 Mooncake。MoRI ↗ 是著重 RDMA 與 GPU integration 的 high-performance communication interface,可用於 network collective operation、expert parallel kernel 等;最近加入 PyTorch ecosystem 的 Mooncake ↗,則支援 prefill-decode disaggregation 與多種 fault-tolerant multi-node 功能。

Atomic Claim 124/574 · 2026-02-16_inferencex-v2-nvidia-blackwell-vs::IX2-0126

Claim: NVIDIAdisaggregated inference 配置使用 Dynamo。
Frame: RELATION · Mode: ASSERTED · Mapping: COMPLETE
開啟逐條審核

Atomic Claim 125/574 · 2026-02-16_inferencex-v2-nvidia-blackwell-vs::IX2-0130

Claim: MoRI 是聚焦 RDMAGPU 整合的高效能 communication interface,可用於 network collective operations 與 expert parallel kernels 等應用。
Frame: NARY_RELATION · Mode: ASSERTED · Mapping: COMPLETE
開啟逐條審核

Atomic Claim 126/574 · 2026-02-16_inferencex-v2-nvidia-blackwell-vs::IX2-0127

Claim: Dynamo 是為 multi-node distributed inference 設計的 inference framework,支援 prefill-decode disaggregation、request routing 與 KV cache offloading 等技術。
Frame: NARY_RELATION · Mode: ASSERTED · Mapping: COMPLETE
開啟逐條審核

Atomic Claim 127/574 · 2026-02-16_inferencex-v2-nvidia-blackwell-vs::IX2-0131

Claim: Mooncake 最近加入 PyTorch ecosystem,支援 prefill-decode disaggregation,以及多種具 fault tolerant 能力的 multi-node features。
Frame: NARY_RELATION · Mode: ASSERTED · Mapping: COMPLETE
開啟逐條審核

Atomic Claim 128/574 · 2026-02-16_inferencex-v2-nvidia-blackwell-vs::IX2-0128

Claim: Dynamo 與 inference engine 無關,因此 benchmark 可分別使用 SGLang 與 TRT LLM 作為 backend。
Frame: RELATION · Mode: ASSERTED · Mapping: COMPLETE
開啟逐條審核

Atomic Claim 129/574 · 2026-02-16_inferencex-v2-nvidia-blackwell-vs::IX2-0129

Claim:AMD,InferenceX 使用 SGLang 搭配兩種不同的 KV cache transfer frameworks:MoRIMooncake
Frame: NARY_RELATION · Mode: ASSERTED · Mapping: PARTIAL
開啟逐條審核

DeepSeek Disagg +WideEP Results Deep Dive

At almost all interactivity levels, disagg outperform aggregated inference (grey lines) in terms of total token throughput per GPU. Multi-node disaggregrated prefill framemogs single node aggregrated serving.

幾乎所有 interactivity level 下,disagg 的 total token throughput per GPU 都優於 aggregated inference(灰線)。Multi-node disaggregated prefill 直接 framemog single-node aggregated serving。

Atomic Claim 130/574 · 2026-02-16_inferencex-v2-nvidia-blackwell-vs::IX2-0132

Claim: 在幾乎所有 interactivity 水準下,disagg 的每 GPU 總 token throughput 都優於 aggregated inference(灰線)。
Frame: COMPARISON · Mode: ASSERTED · Mapping: PARTIAL
開啟逐條審核

Atomic Claim 131/574 · 2026-02-16_inferencex-v2-nvidia-blackwell-vs::IX2-0133

Claim: Multi-node disaggregated prefill 明顯優於 single-node aggregated serving。
Frame: COMPARISON · Mode: ASSERTED · Mapping: PARTIAL
開啟逐條審核

image

Source: SemiAnalysis InferenceX

Nvidia continues to push new updates for B200/GB200 FP8. The latest data on DeepSeek FP8 B200 TRT single node (both MTP enabled/disabled) vs GB200 Dynamo+TRT disagg (both MTP enabled/disabled). This indicates consistent engineering effort to improve rack-scale inference software and wideEP kernels.

Nvidia 持續替 B200/GB200 FP8 推新 update。最新資料比較 DeepSeek FP8 B200 TRT single-node(MTP 開/關)與 GB200 Dynamo+TRT disagg(MTP 開/關),可以看出 Nvidia 一直投入 engineering effort 改善 rack-scale inference software 與 wideEP kernel。

Atomic Claim 132/574 · 2026-02-16_inferencex-v2-nvidia-blackwell-vs::IX2-0134

Claim: Nvidia 持續針對 B200GB200 FP8 推出新更新。
Frame: NARY_RELATION · Mode: ASSERTED · Mapping: COMPLETE
開啟逐條審核

Atomic Claim 133/574 · 2026-02-16_inferencex-v2-nvidia-blackwell-vs::IX2-0135

Claim: 最新資料比較 DeepSeek FP8 B200 TRT single-node(MTP 開啟/關閉)與 GB200 Dynamo+TRT disagg(MTP 開啟/關閉)。
Frame: COMPARISON · Mode: ASSERTED · Mapping: COMPLETE
開啟逐條審核

image

Source: SemiAnalysis InferenceX

When comparing MI355X disaggregated inference vs aggregated inference, we noticed a similar pattern. Disaggregated inference only overtakes aggregated inference at low interactivity, high batch sizes. This is true across FP4, and it is likely due to poorly optimized kernels.

比較 MI355X disaggregated inference 與 aggregated inference 時,我們看到類似現象:只有在 low interactivity、high batch size 區域,disaggregated inference 才真正超越 aggregated inference。FP4 也一樣,可能原因是 kernel optimization 還不夠成熟。

Atomic Claim 134/574 · 2026-02-16_inferencex-v2-nvidia-blackwell-vs::IX2-0137

Claim: 比較 MI355X disaggregated inference 與 aggregated inference 時,也觀察到類似模式。
Frame: ATTRIBUTE · Mode: ASSERTED · Mapping: PARTIAL
開啟逐條審核

Atomic Claim 135/574 · 2026-02-16_inferencex-v2-nvidia-blackwell-vs::IX2-0138

Claim: Disaggregated inference 只有在低 interactivity、高 batch size 時才會超越 aggregated inference。
Frame: ATTRIBUTE · Mode: ASSERTED · Mapping: COMPLETE
開啟逐條審核

Atomic Claim 136/574 · 2026-02-16_inferencex-v2-nvidia-blackwell-vs::IX2-0139

Claim: 這個現象在 FP4 也成立。
Frame: ATTRIBUTE · Mode: ASSERTED · Mapping: COMPLETE
開啟逐條審核

Atomic Claim 137/574 · 2026-02-16_inferencex-v2-nvidia-blackwell-vs::IX2-0140

Claim: 關於 FP4:SemiAnalysis 認為原因很可能是 kernels 最佳化不足。
Frame: ATTRIBUTE · Mode: INFERRED · Mapping: PARTIAL
開啟逐條審核

image

Source: SemiAnalysis InferenceX

When composing disagg prefill+wideEP with FP4 on the MI355X, we observe suffers subpar performance.

MI355X 把 disagg prefill+wideEP 與 FP4 疊在一起時,performance 明顯低於應有水準。

Atomic Claim 138/574 · 2026-02-16_inferencex-v2-nvidia-blackwell-vs::IX2-0141

Claim:MI355X 上同時組合 disagg prefill+wideEP 與 FP4 時,實際效能明顯不佳。
Frame: ATTRIBUTE · Mode: ASSERTED · Mapping: COMPLETE
開啟逐條審核

Although theoretical modeling shows that disagg inference on MI355Xs should perform way better than single node, disagg actually performs worse for higher interactivity levels due to a lack of kernel and collective optimization in the ROCm software stack when composing multiple SOTA inference optimizations together.

理論模型顯示,MI355X 的 disagg inference 應該大幅優於 single-node;但在較高 interactivity 下,實際 disagg 反而更慢,原因是 ROCm software stack 在把多項 SOTA inference optimization 組合起來時,kernel 與 collective optimization 還不足。

Atomic Claim 139/574 · 2026-02-16_inferencex-v2-nvidia-blackwell-vs::IX2-0142

Claim: 雖然理論模型顯示 MI355X 的 disagg inference 應明顯優於 single-node,但在較高 interactivity 下反而更慢,原因是 ROCm software stack 在同時組合多種 SOTA inference optimizations together 時,缺乏足夠的 kernel 與 collective 最佳化。
Frame: COMPARISON · Mode: ASSERTED · Mapping: COMPLETE
開啟逐條審核

image

Source: SemiAnalysis InferenceX

Nvidia TensorRT LLM and NVL72

TensorRT LLM already serves billions of tokens per hour globally across providers like TogetherAI and other advanced providers, and it has really allowed the GB200 NVL72 and GB300 NVL72 to shine, delivering more than double the performance at high throughput. MTP boosts these results even further, making use of the chips’ full potential.

TensorRT-LLM 現在已經在 TogetherAI 與其他 advanced provider 的 production 中,每小時服務數十億 token。它真的讓 GB200 NVL72、GB300 NVL72 的能力完整發揮,在 high-throughput 區域可帶來超過 2x performance;MTP 又能把結果進一步往上推,更接近 chip 的完整潛力。

Atomic Claim 140/574 · 2026-02-16_inferencex-v2-nvidia-blackwell-vs::IX2-0144

Claim: TensorRT LLM 已充分發揮 GB200 NVL72GB300 NVL72 的能力,在高 throughput 區間提供超過兩倍效能。
Frame: COMPARISON · Mode: ASSERTED · Mapping: COMPLETE
開啟逐條審核

Atomic Claim 141/574 · 2026-02-16_inferencex-v2-nvidia-blackwell-vs::IX2-0143

Claim: TensorRT LLM 已在 TogetherAI 等全球 providers 上,每小時服務數十億 tokens。
Frame: ATTRIBUTE · Mode: ASSERTED · Mapping: PARTIAL
開啟逐條審核

Atomic Claim 142/574 · 2026-02-16_inferencex-v2-nvidia-blackwell-vs::IX2-0145

Claim: MTP 能進一步提高這些結果,更充分利用晶片潛力。
Frame: ATTRIBUTE · Mode: ASSERTED · Mapping: PARTIAL
開啟逐條審核

image

image

Source: SemiAnalysis InferenceX

The benefits delivered from the larger world size of the NVL72 family is also evident if we look at cost graphs. At a fixed interactivity level of 60 tok/s/user, each GB200 NVL GPU produces slightly less than triple the number of tokens/s than each B200 does.

NVL72 family 較大 world size 的好處,在 cost graph 也很明顯。固定 interactivity = 60 tok/s/user 時,每顆 GB200 NVL GPU 產生的 tokens/s 接近 B200 的三倍。

Atomic Claim 143/574 · 2026-02-16_inferencex-v2-nvidia-blackwell-vs::IX2-0147

Claim: 在固定 interactivity 60 tok/s/user 下,每顆 GB200 NVL GPU 產生的 tokens/s 接近 B200 的三倍。
Frame: COMPARISON · Mode: ASSERTED · Mapping: COMPLETE
開啟逐條審核

SemiAnalysis InferenceX is free open source software and reader-supported. To receive new posts and support our work, consider becoming a free or paid subscriber.

SemiAnalysis InferenceX 是免費 open-source software,也由讀者支持。如果想收到新文章並支持我們,可以考慮成為免費或付費訂閱者。

Subscribed

image

Source: SemiAnalysis InferenceX

This gap shrinks as interactivity increases. At 130 tok/s/user, the GB200 NVL72 has nearly no advantage and is even more expensive on a $/Million tokens basis. At low batch sizes, the inference workload shrinks enough to fit within a single HGX node’s NVLink domain (i.e. 8 GPUs), and the GB200 NVL72’s larger scale-out advantage starts to disappear.

但 interactivity 越高,這個差距就越小。到了 130 tok/s/user,GB200 NVL72 幾乎沒有優勢,甚至以 $/Million tokens 衡量還更貴。Low batch size 時,inference workload 已縮小到可以完全塞進單一 HGX node 的 NVLink domain(8 GPUs),因此 GB200 NVL72 較大 scale-up domain 的優勢開始消失。

Atomic Claim 144/574 · 2026-02-16_inferencex-v2-nvidia-blackwell-vs::IX2-0149

Claim: 在 130 tok/s/user 時,GB200 NVL72 幾乎已沒有優勢。
Frame: ATTRIBUTE · Mode: ASSERTED · Mapping: COMPLETE
開啟逐條審核

Atomic Claim 145/574 · 2026-02-16_inferencex-v2-nvidia-blackwell-vs::IX2-0150

Claim: 在 130 tok/s/user 時,GB200 NVL72 以每百萬 token 成本計算甚至更貴。
Frame: ATTRIBUTE · Mode: ASSERTED · Mapping: COMPLETE
開啟逐條審核

Atomic Claim 146/574 · 2026-02-16_inferencex-v2-nvidia-blackwell-vs::IX2-0151

Claim: 在低 batch size 下,推論 workload 會縮小到足以放進單一 HGX node 的 NVLink domain,也就是 8 顆 GPUs
Frame: ATTRIBUTE · Mode: ASSERTED · Mapping: COMPLETE
開啟逐條審核

Atomic Claim 147/574 · 2026-02-16_inferencex-v2-nvidia-blackwell-vs::IX2-0152

Claim: 因此 GB200 NVL72 較大 scale-out domain 的優勢開始消失。
Frame: COMPARISON · Mode: ASSERTED · Mapping: PARTIAL
開啟逐條審核

image

Source: SemiAnalysis InferenceX

Nvidia versus AMD Disagg Prefill

With today’s release of InferenceXv2, for the first time the ML community is able to see a full Pareto frontier for open-source MI355X distributed inference. We show Pareto curves for the B200 and MI355X with and without enabling MTP.

隨 InferenceXv2 今天發布,ML community 第一次能看到 open-source MI355X distributed inference 的完整 Pareto frontier。我們會分別展示 B200、MI355X 在開啟與未開啟 MTP 時的 Pareto curve。

Atomic Claim 148/574 · 2026-02-16_inferencex-v2-nvidia-blackwell-vs::IX2-0153

Claim: 隨著 InferenceXv2 發布,ML 社群首次能看到開源 MI355X distributed inference 的完整 Pareto frontier。
Frame: RELATION · Mode: ASSERTED · Mapping: COMPLETE
開啟逐條審核

Atomic Claim 149/574 · 2026-02-16_inferencex-v2-nvidia-blackwell-vs::IX2-0154

Claim: 文章展示 B200MI355X 在開啟與關閉 MTP 兩種情況下的 Pareto curves。
Frame: ATTRIBUTE · Mode: ASSERTED · Mapping: PARTIAL
開啟逐條審核

For FP8 disagg prefill, MI355X (MoRI SGLang) is quite competitive with B200 (Dynamo SGLang). Wide EP is not used for either of these configs as all prefill/decode instances run using EP8 at the most. At both ends of the throughput versus interactivity Pareto frontier, MI355X falls behind the B200 slightly. However, MI355X disagg has a slight advantage for certain levels of interactivity in the middle of the curve. Both the B200 and the MI355X benefit from employing MTP, and we observe the same relative performance improvement for both chips when using MTP.

FP8 disagg prefill 方面,MI355X(MoRI SGLang)與 B200(Dynamo SGLang)其實相當接近。這兩種 configuration 都沒有使用 wide EP,因為 prefill/decode instance 最多只跑到 EP8。在 throughput-vs-interactivity Pareto frontier 兩端,MI355X 都略輸 B200;但 curve 中段某些 interactivity,MI355X disagg 反而有一點優勢。B200、MI355X 使用 MTP 都會受益,而且兩顆 chip 的相對 performance improvement 幅度很接近。

Atomic Claim 150/574 · 2026-02-16_inferencex-v2-nvidia-blackwell-vs::IX2-0155

Claim:FP8 disagg prefill 而言,MI355XMoRI SGLang)與 B200(Dynamo SGLang)的競爭力相當接近。
Frame: ATTRIBUTE · Mode: ASSERTED · Mapping: COMPLETE
開啟逐條審核

Atomic Claim 151/574 · 2026-02-16_inferencex-v2-nvidia-blackwell-vs::IX2-0156

Claim: 這兩個 configurations 都沒有使用 wide EP,因為所有 prefilldecode instances 最多只使用 EP8。
Frame: RELATION · Mode: ASSERTED · Mapping: COMPLETE
開啟逐條審核

Atomic Claim 152/574 · 2026-02-16_inferencex-v2-nvidia-blackwell-vs::IX2-0157

Claim: 在 throughput vs interactivity Pareto frontier 的兩端,MI355X 都略微落後 B200
Frame: COMPARISON · Mode: ASSERTED · Mapping: COMPLETE
開啟逐條審核

Atomic Claim 153/574 · 2026-02-16_inferencex-v2-nvidia-blackwell-vs::IX2-0158

Claim: 但在曲線中段某些 interactivity 水準,MI355X disagg 反而有小幅優勢。
Frame: ATTRIBUTE · Mode: ASSERTED · Mapping: COMPLETE
開啟逐條審核

Atomic Claim 154/574 · 2026-02-16_inferencex-v2-nvidia-blackwell-vs::IX2-0159

Claim: B200MI355X 都能受惠於 MTP,兩顆晶片使用 MTP 時的相對效能改善幅度相近。
Frame: RELATION · Mode: ASSERTED · Mapping: COMPLETE
開啟逐條審核

image

Source: SemiAnalysis InferenceX

However, if we were to only measure output (decode) token throughput, we see that output token throughput is much higher for the B200 than for the MI355X at lower interactivity levels. Note that when looking at output token only throughput for disaggregated inference configurations, we normalize throughout by the number of decode GPUs, not total GPUs. It is possible that different numbers of GPUs are used for output when running inference jobs on the B200 and MI355X, but the bottom line is that whatever configuration decode is run on, B200 gets the decode job done faster.

不過若只看 output(decode)token throughput,在較低 interactivity 區域,B200 明顯高於 MI355X。要注意,disaggregated inference configuration 的 output-token-only throughput,是用 decode GPU 數量做 normalization,而不是 total GPU 數。B200、MI355X inference job 可能用了不同數量 GPU 來產生 output,但結論不變:無論 decode configuration 怎麼配置,B200 都能更快完成 decode。

Atomic Claim 155/574 · 2026-02-16_inferencex-v2-nvidia-blackwell-vs::IX2-0160

Claim: 但若只衡量 output(decode)token throughput,在較低 interactivity 下,B200 明顯高於 MI355X
Frame: COMPARISON · Mode: ASSERTED · Mapping: COMPLETE
開啟逐條審核

Atomic Claim 156/574 · 2026-02-16_inferencex-v2-nvidia-blackwell-vs::IX2-0161

Claim: 分析 disaggregated inference 的 output-token-only throughput 時,InferenceX 是以 decode GPUs 數量做正規化,而非全部 GPUs 數量。
Frame: ATTRIBUTE · Mode: ASSERTED · Mapping: PARTIAL
開啟逐條審核

Atomic Claim 157/574 · 2026-02-16_inferencex-v2-nvidia-blackwell-vs::IX2-0162

Claim: B200MI355X 執行 inference job 時,可能使用不同數量的 GPUs 負責 output。
Frame: NARY_RELATION · Mode: ASSERTED · Mapping: PARTIAL
開啟逐條審核

Atomic Claim 158/574 · 2026-02-16_inferencex-v2-nvidia-blackwell-vs::IX2-0163

Claim: 但無論 decode 採哪種配置,B200 都能更快完成 decode 工作。
Frame: RELATION · Mode: ASSERTED · Mapping: COMPLETE
開啟逐條審核

SemiAnalysis is free open source software and reader-supported. To receive new posts and support our work, consider becoming a free or paid subscriber.

SemiAnalysis 是免費 open-source software,也由讀者支持。如果想收到新文章並支持我們,可以考慮成為免費或付費訂閱者。

Atomic Claim 159/574 · 2026-02-16_inferencex-v2-nvidia-blackwell-vs::IX2-0164

Claim: SemiAnalysis 的相關軟體為免費開源,並由讀者支持。
Frame: CLAIM_ONLY · Mode: ASSERTED · Mapping: CLAIM_ONLY
開啟逐條審核

Subscribed

image

Source: SemiAnalysis InferenceX

Despite the MI355X being competitive in FP8 disagg, its FP4 performance suffers from composability issues. AMD single node FP4 performance is decent, but when we compare AMD FP4 disagg prefill to Nvidia, performance is subpar and the MI355X gets absolutely mogged by Nvidia’s B200. In a 1k1k scenario, the MI355X (MoRI SGLang) with MTP barely manages to beat the B200 (Dynamo SGLang) without MTP.

雖然 MI355X 在 FP8 disagg 很有競爭力,但 FP4 performance 受到 composability 問題拖累。AMD single-node FP4 表現還算不錯,可是把 AMD FP4 disagg prefill 和 Nvidia 比,performance 明顯偏弱,MI355X 被 B200 徹底 mogged。在 1k1k scenario 中,開 MTP 的 MI355X(MoRI SGLang)也只是勉強打贏沒開 MTP 的 B200(Dynamo SGLang)。

Atomic Claim 160/574 · 2026-02-16_inferencex-v2-nvidia-blackwell-vs::IX2-0165

Claim: 雖然 MI355XFP8 disagg 下具競爭力,但 FP4 效能受到 composability 問題拖累。
Frame: ATTRIBUTE · Mode: ASSERTED · Mapping: COMPLETE
開啟逐條審核

Atomic Claim 161/574 · 2026-02-16_inferencex-v2-nvidia-blackwell-vs::IX2-0166

Claim: AMD 的 single-node FP4 效能本身不差。
Frame: ATTRIBUTE · Mode: ASSERTED · Mapping: PARTIAL
開啟逐條審核

Atomic Claim 162/574 · 2026-02-16_inferencex-v2-nvidia-blackwell-vs::IX2-0167

Claim: 但比較 AMD FP4 disagg prefillNvidia 時,MI355X 表現明顯不佳,遠落後 NvidiaB200
Frame: ATTRIBUTE · Mode: ASSERTED · Mapping: COMPLETE
開啟逐條審核

Atomic Claim 163/574 · 2026-02-16_inferencex-v2-nvidia-blackwell-vs::IX2-0168

Claim: 在 1k1k 情境下,開啟 MTPMI355XMoRI SGLang)僅勉強超越未開啟 MTPB200(Dynamo SGLang)。
Frame: ATTRIBUTE · Mode: ASSERTED · Mapping: PARTIAL
開啟逐條審核

image

Source: SemiAnalysis InferenceX

Once we bring Dynamo TRT-LLM into the equation, the B200’s performance is boosted even more to the point that the MI355X even with MTP can’t match the B200’s performance with Dynamo TRT-LLM and MTP. The MI355X can only match the B200 (without MTP) in performance by using MTP, and only for a range of interactivities from ~60 tok/s/user through ~120 tok/s/user.

若再把 Dynamo TRT-LLM 放進比較,B200 performance 還會進一步拉高;即使 MI355X 開 MTP,也追不上同樣使用 Dynamo TRT-LLM + MTP 的 B200。MI355X 必須靠 MTP,才有機會追平『沒開 MTP』的 B200,而且只限大約 60–120 tok/s/user 這段 interactivity。

Atomic Claim 164/574 · 2026-02-16_inferencex-v2-nvidia-blackwell-vs::IX2-0169

Claim: 一旦加入 Dynamo TRT-LLMB200 效能進一步提高,使即使開啟 MTPMI355X 也無法匹配搭配 Dynamo TRT-LLMMTPB200
Frame: ATTRIBUTE · Mode: ASSERTED · Mapping: PARTIAL
開啟逐條審核

Atomic Claim 165/574 · 2026-02-16_inferencex-v2-nvidia-blackwell-vs::IX2-0170

Claim: MI355X 只有在開啟 MTP 後,且約 60~120 tok/s/user 的 interactivity 區間內,才能匹配未開啟 MTPB200
Frame: ATTRIBUTE · Mode: ASSERTED · Mapping: COMPLETE
開啟逐條審核

image

Source: SemiAnalysis InferenceX

When comparing Dynamo TRTLLM B200 disagg prefill to SGLang MoRI MI355 disagg prefill, AMD gets framemogged due to the more mature implementation of disagg prefill on TRTLLM.

比較 Dynamo TRTLLM B200 disagg prefill 與 SGLang MoRI MI355X disagg prefill 時,因為 TRTLLM 的 disagg prefill implementation 更成熟,AMD 直接被 framemog。

Atomic Claim 166/574 · 2026-02-16_inferencex-v2-nvidia-blackwell-vs::IX2-0171

Claim: 比較 Dynamo TRTLLM B200 disagg prefillSGLang MoRI MI355 disagg prefillAMD 明顯落後,主要因為 TRTLLM 的 disagg prefill implementation 更成熟。
Frame: ATTRIBUTE · Mode: ASSERTED · Mapping: COMPLETE
開啟逐條審核

image

Source: SemiAnalysis InferenceX

image

Source: Dwarkesh Podcast and SemiAnalysis

The diagram below shows us the various parallelism configurations that form up the MI355X (MoRI SGLang) Pareto frontier. Note that currently, wide EP is not employed for any points (i.e., configurations with EP 16, 32, etc.).

下圖列出構成 MI355X(MoRI SGLang)Pareto frontier 的各種 parallelism configuration。要注意,目前沒有任何 point 使用 wide EP,也就是還沒有 EP16、EP32 等 configuration。

Atomic Claim 167/574 · 2026-02-16_inferencex-v2-nvidia-blackwell-vs::IX2-0172

Claim: 下方圖表展示構成 MI355XMoRI SGLang)Pareto frontier 的各種 parallelism configurations。
Frame: ATTRIBUTE · Mode: ASSERTED · Mapping: PARTIAL
開啟逐條審核

Atomic Claim 168/574 · 2026-02-16_inferencex-v2-nvidia-blackwell-vs::IX2-0173

Claim: 目前所有點都沒有使用 wide EP,也就是沒有 EP16、EP32 等 configurations。
Frame: ATTRIBUTE · Mode: ASSERTED · Mapping: PARTIAL
開啟逐條審核

image

Source: SemiAnalysis InferenceX

Unpacking Inference Providers’ Unit Economics

Below is a list on OpenRouter of all inference providers that serve DeepSeek R1 0528 FP8 along with their cost per million input/output tokens and average interactivity listed on. Disregarding Chutes, the middle of the pack provider serves at an interactivity of around 35 tok/s/user.

下方列出 OpenRouter 上所有提供 DeepSeek R1 0528 FP8 的 inference provider,包括每百萬 input/output token 價格與平均 interactivity。若排除 Chutes,位於中間水準的 provider 大約提供 35 tok/s/user。

Atomic Claim 169/574 · 2026-02-16_inferencex-v2-nvidia-blackwell-vs::IX2-0174

Claim: 文章列出 OpenRouter 上所有提供 DeepSeek R1 0528 FP8 的 inference providers,以及各自每百萬 input/output tokens 價格與平均 interactivity。
Frame: NARY_RELATION · Mode: ASSERTED · Mapping: PARTIAL
開啟逐條審核

image

Source: OpenRouter

We can then use real InferenceX data to interpolate the cost per million input/output tokens at an interactivity level of 35 tok/sec/user, which is a reasonable interactivity level given the data above.

接著可以用真實 InferenceX data,在 35 tok/s/user 的 interactivity 水準插值估算每百萬 input/output token 的成本;從上面的市場資料來看,35 tok/s/user 是很合理的代表值。

Atomic Claim 170/574 · 2026-02-16_inferencex-v2-nvidia-blackwell-vs::IX2-0176

Claim: 因此可以利用真實 InferenceX 資料,插值估算在 35 tok/sec/user interactivity 下每百萬 input/output tokens 的成本;根據上述資料,這是一個合理 interactivity 水準。
Frame: ATTRIBUTE · Mode: ASSERTED · Mapping: PARTIAL
開啟逐條審核

As we mention later in the article, this is best understood as _baseline _data and not completely representative of real-world inference, mainly because InferenceX benchmarks on random data and disables prefix caching. In other words, performance/cost will be _at least _this good. It is also important to note that there are not data points for each GPU at _each _interactivity level. Thus we cannot make _exact _comparisons at each degree of interactivity. We nevertheless think the bar chart comparisons presented below are (very) reasonable interpolations in lieu of using exact data points.

如本文後面會提到,這些數字最好理解成 baseline,而不是完全等同 real-world inference。主要原因是 InferenceX 使用 random data 做 benchmark,並關閉 prefix caching;換句話說,真實 performance/cost 至少應該可以做到這麼好。另外,每一顆 GPU 並不是在每一個 interactivity level 都有實測 data point,因此不能做完全精確的一對一比較。即便如此,在缺少 exact data point 的情況下,我們仍認為下面 bar chart 的 interpolation 是相當合理的。

Atomic Claim 171/574 · 2026-02-16_inferencex-v2-nvidia-blackwell-vs::IX2-0177

Claim: 如文章後段說明,這些結果更適合視為 baseline,而不完全代表 real-world inference,主要因為 InferenceX 使用隨機資料做 benchmark,且關閉 prefix caching
Frame: RELATION · Mode: ASSERTED · Mapping: COMPLETE
開啟逐條審核

Atomic Claim 172/574 · 2026-02-16_inferencex-v2-nvidia-blackwell-vs::IX2-0178

Claim: 關於 InferenceX:換句話說,實際 production 的 performance/cost 至少應該不會比這個 baseline 更差。
Frame: ATTRIBUTE · Mode: ASSERTED · Mapping: COMPLETE
開啟逐條審核

Atomic Claim 173/574 · 2026-02-16_inferencex-v2-nvidia-blackwell-vs::IX2-0179

Claim: 關於 GPU:另一項需要注意的是,不是每顆 GPU 在每個 interactivity 水準都有實際 data point。
Frame: ATTRIBUTE · Mode: ASSERTED · Mapping: COMPLETE
開啟逐條審核

Atomic Claim 174/574 · 2026-02-16_inferencex-v2-nvidia-blackwell-vs::IX2-0180

Claim: 關於 InferenceX:因此無法在每個 interactivity 水準都做完全精確的一對一比較。
Frame: ATTRIBUTE · Mode: ASSERTED · Mapping: COMPLETE
開啟逐條審核

Atomic Claim 175/574 · 2026-02-16_inferencex-v2-nvidia-blackwell-vs::IX2-0181

Claim: 儘管如此,SemiAnalysis 認為下方 bar-chart comparisons 在缺少完整 data points 的情況下,仍屬非常合理的插值。
Frame: CLAIM_ONLY · Mode: ASSERTED · Mapping: CLAIM_ONLY
開啟逐條審核

Comparing disagg+wideEP configs at this interactivity level, we see just how effective distributed inference techniques are when it comes to both perf/TCO and overall throughput. We also see how large scale up domains (like GB300 and GB200 NVL72) absolutely dominate in total throughput per GPU.

在這個 interactivity level 比較 disagg+wideEP configuration,可以清楚看到 distributed inference technique 對 perf/TCO 與 overall throughput 有多有效;同時也能看到 GB300、GB200 NVL72 這種 large scale-up domain,在 total throughput per GPU 上具有壓倒性優勢。

Atomic Claim 176/574 · 2026-02-16_inferencex-v2-nvidia-blackwell-vs::IX2-0183

Claim: 也可以看到像 GB300GB200 NVL72 這類大型 scale-up domains,在每 GPU 總 throughput 上具有明顯優勢。
Frame: ATTRIBUTE · Mode: ASSERTED · Mapping: COMPLETE
開啟逐條審核

It is interesting to note that at this interactivity level (on an 8k1k workload type), the B200 can achieve the best perf/TCO when MTP is enabled. Below we also list the Total Cost of Ownership (TCO) (Owning – Hyperscaler) for each GPU:

有趣的是,在這個 interactivity level、8k1k workload 下,開啟 MTP 的 B200 可以做到最佳 perf/TCO。下方我們也列出每款 GPU 的 Total Cost of Ownership(TCO,Owning-Hyperscaler)。

Atomic Claim 177/574 · 2026-02-16_inferencex-v2-nvidia-blackwell-vs::IX2-0184

Claim: 在此 interactivity 水準(8k1k workload)下,開啟 MTPB200 可以達到最佳 perf/TCO。
Frame: ATTRIBUTE · Mode: ASSERTED · Mapping: PARTIAL
開啟逐條審核

Atomic Claim 178/574 · 2026-02-16_inferencex-v2-nvidia-blackwell-vs::IX2-0185

Claim: 文章也列出各 GPU 的 Total Cost of Ownership(TCO,Owning – Hyperscaler)。
Frame: ATTRIBUTE · Mode: ASSERTED · Mapping: COMPLETE
開啟逐條審核

image

Source: SemiAnalysis TCO Model

image

image

image

Source: SemiAnalysis InferenceX

Let’s use the findings above to dig deeper into the unit economics of serving LLMs at scale. From the OpenRouter data above, we see that Crusoe serves at 36 tok/sec/user at 5.40/M output tokens. If we assume no cache hits and that Crusoe is using at least H200s with SOTA inference techniques like MTP, disagg, and wide EP, the data above suggests they incur a cost of _no more than _/M input tokens and $2.955/M output tokens for a profit margin of up to 83% gross margin (depreciation counted in cost of goods sold) on input tokens and 45% gross margin on output tokens.

接著把上面的結果拿來深入看大規模 LLM serving 的 unit economics。從 OpenRouter 資料可看到,Crusoe 以 36 tok/sec/user 提供服務,價格是每百萬 input token $1.35、每百萬 output token $5.40。假設沒有 cache hit,而且 Crusoe 至少使用 H200,並搭配 MTP、disagg、wide EP 等 SOTA inference technique,上面的資料顯示其成本最多約為每百萬 input token $0.226、每百萬 output token $2.955。這相當於 input token gross margin 最高約 83%、output token gross margin 約 45%(depreciation 已計入 COGS)。

Atomic Claim 179/574 · 2026-02-16_inferencex-v2-nvidia-blackwell-vs::IX2-0189

Claim: 假設 Crusoe 無 cache hit、至少使用 H200,並採 MTP、disaggregated serving、wide EP 等 SOTA inference techniques,SemiAnalysis 估計 input token cost 不超過 2.955/M;對應 gross margin 分別最高約 83% 與 45%,折舊計入 COGS。
Frame: ATTRIBUTE · Mode: HYPOTHETICAL · Mapping: PARTIAL
開啟逐條審核

Atomic Claim 180/574 · 2026-02-16_inferencex-v2-nvidia-blackwell-vs::IX2-0186

Claim: 接著利用上述結果進一步分析大規模 serving LLMs 的 unit economics。
Frame: ATTRIBUTE · Mode: ASSERTED · Mapping: COMPLETE
開啟逐條審核

Atomic Claim 181/574 · 2026-02-16_inferencex-v2-nvidia-blackwell-vs::IX2-0187

Claim: 根據 OpenRouter 資料,Crusoe 以 36 tok/sec/user 提供服務,input token 價格為每百萬 1.35 美元,output token 為每百萬 5.40 美元。
Frame: RELATION · Mode: ASSERTED · Mapping: COMPLETE
開啟逐條審核

Atomic Claim 182/574 · 2026-02-16_inferencex-v2-nvidia-blackwell-vs::IX2-0188

Claim: 若假設沒有 cache hits,且 Crusoe 至少使用 H200s,並搭配 MTP、disagg 等 SOTA inference techniques。
Frame: RELATION · Mode: HYPOTHETICAL · Mapping: COMPLETE
開啟逐條審核

SemiAnalysis InferenceX is free open source software and reader-supported. To receive new posts and support our work, consider becoming a free or paid subscriber.

SemiAnalysis InferenceX 是免費 open-source software,也由讀者支持。如果想收到新文章並支持我們,可以考慮成為免費或付費訂閱者。

Subscribed

Of course, these assumptions may not be _exactly _correct and these calculations don’t account for downtime or underutilization, but this gives an idea of some cool math you can do with InferenceX data. More analysis on the economics of inference can be found in the SemiAnalysis Tokenomics Model .

當然,上述 assumption 不一定完全精確,計算也沒有納入 downtime 或 underutilization;但它可以展示 InferenceX data 能做出哪些有意思的 economics analysis。更多 inference economics 分析可參考 SemiAnalysis Tokenomics Model ↗。

Atomic Claim 183/574 · 2026-02-16_inferencex-v2-nvidia-blackwell-vs::IX2-0190

Claim: 關於 InferenceX:當然,這些假設可能不完全正確,而且計算沒有納入 downtime 或 underutilization。
Frame: ATTRIBUTE · Mode: HYPOTHETICAL · Mapping: COMPLETE
開啟逐條審核

Atomic Claim 184/574 · 2026-02-16_inferencex-v2-nvidia-blackwell-vs::IX2-0191

Claim: 但這可以示範如何利用 InferenceX 資料進行有價值的 unit-economics 計算。
Frame: ATTRIBUTE · Mode: ASSERTED · Mapping: COMPLETE
開啟逐條審核

Atomic Claim 185/574 · 2026-02-16_inferencex-v2-nvidia-blackwell-vs::IX2-0192

Claim: 更多推論經濟性分析可參考 SemiAnalysis Tokenomics Model
Frame: ATTRIBUTE · Mode: ASSERTED · Mapping: PARTIAL
開啟逐條審核

The OpenRouter data also shows Nebius AI Studio (Fast) serving DeepSeek FP4 at 167 tok/sec/user at 6/M output tokens. Adjusting the interactivity level in InferenceX accordingly and we see the following data.

OpenRouter data 也顯示 Nebius AI Studio(Fast)以 167 tok/sec/user 提供 DeepSeek FP4,價格為每百萬 input token $2、output token $6。把 InferenceX 的 interactivity 調到相同水準後,可以得到下面的比較結果。

Atomic Claim 186/574 · 2026-02-16_inferencex-v2-nvidia-blackwell-vs::IX2-0193

Claim: OpenRouter 資料也顯示,Nebius AI Studio(Fast)以 167 tok/sec/user 提供 DeepSeek FP4,input token 每百萬 2 美元、output token 每百萬 6 美元。
Frame: NARY_RELATION · Mode: ASSERTED · Mapping: COMPLETE
開啟逐條審核

Atomic Claim 187/574 · 2026-02-16_inferencex-v2-nvidia-blackwell-vs::IX2-0194

Claim:InferenceX 的 interactivity 調整到對應水準後,可以得到後續資料。
Frame: ATTRIBUTE · Mode: ASSERTED · Mapping: COMPLETE
開啟逐條審核

image

image

image

Source: SemiAnalysis InferenceX

At this high of interactivity, it becomes necessary to employ speculative decoding techniques like MTP to achieve high enough throughput to make inference economical. Luckily, MTP can increase throughput with relatively low risk to overall model accuracy. We will go on to talk more about MTP, and how it can be applied to increase throughput / decrease cost, in later sections of this article.

在這麼高的 interactivity 下,要維持足夠 throughput、讓 inference economics 成立,就幾乎必須使用 MTP 這類 speculative decoding technique。好消息是,MTP 能在對 model accuracy 風險相對低的情況下提高 throughput。本文後面會再深入討論 MTP,以及如何用它增加 throughput、降低 cost。

Atomic Claim 188/574 · 2026-02-16_inferencex-v2-nvidia-blackwell-vs::IX2-0195

Claim: 在這麼高的 interactivity 下,必須採用 MTPspeculative decoding 技術,才能讓 throughput 高到足以使推論具經濟效益。
Frame: ATTRIBUTE · Mode: ASSERTED · Mapping: COMPLETE
開啟逐條審核

Atomic Claim 189/574 · 2026-02-16_inferencex-v2-nvidia-blackwell-vs::IX2-0196

Claim: MTP 的優點是可提高 throughput,同時對整體模型準確率造成的風險相對低。
Frame: ATTRIBUTE · Mode: ASSERTED · Mapping: COMPLETE
開啟逐條審核

Atomic Claim 190/574 · 2026-02-16_inferencex-v2-nvidia-blackwell-vs::IX2-0197

Claim: 文章後續會進一步討論 MTP
Frame: CLAIM_ONLY · Mode: ASSERTED · Mapping: CLAIM_ONLY
開啟逐條審核

Atomic Claim 191/574 · 2026-02-16_inferencex-v2-nvidia-blackwell-vs::IX2-0198

Claim: 並說明 MTP 如何用來提高 throughput/降低成本。
Frame: ATTRIBUTE · Mode: ASSERTED · Mapping: PARTIAL
開啟逐條審核

Lastly, we show one more chart of an FP8 DeepSeek workload served at 125 tok/s/user. This is another low latency workload where MTP considerably improves economic viability. As with the previous example, we note that at these higher ranges of interactivity, the cheapest configs all use MTP.

最後我們再展示一張 FP8 DeepSeek workload、125 tok/s/user 的圖。這同樣屬於 low-latency workload,而 MTP 可以大幅改善 economic viability。和前一個例子一樣,在這些較高 interactivity 區域,最便宜的 configuration 全部都有使用 MTP。

Atomic Claim 192/574 · 2026-02-16_inferencex-v2-nvidia-blackwell-vs::IX2-0199

Claim: 最後,文章再展示一張 FP8 DeepSeek workload 在 125 tok/s/user 下的圖表。
Frame: ATTRIBUTE · Mode: ASSERTED · Mapping: COMPLETE
開啟逐條審核

Atomic Claim 193/574 · 2026-02-16_inferencex-v2-nvidia-blackwell-vs::IX2-0200

Claim: 這也是一種低延遲 workload,而 MTP 可大幅改善其經濟可行性。
Frame: ATTRIBUTE · Mode: ASSERTED · Mapping: COMPLETE
開啟逐條審核

Atomic Claim 194/574 · 2026-02-16_inferencex-v2-nvidia-blackwell-vs::IX2-0201

Claim: 與前一個例子相同,在這些更高 interactivity 區間,成本最低的 configurations 全都使用 MTP
Frame: COMPARISON · Mode: ASSERTED · Mapping: PARTIAL
開啟逐條審核

image

Source: SemiAnalysis InferenceX

Nvidia Disagg Prefill and WideEP

EP requires all-to-all communication, where every GPU needs to send tokens to every other GPU. This is extremely bandwidth hungry. Recall that Nvidia’s servers have two separate networking domains – the scale-up NVLink domain, and the Scale-out Domain, usually using InfiniBand or Ethernet as the networking protocol.

EP 需要 all-to-all communication,也就是每顆 GPU 都要把 token 傳給其他所有 GPU,對 bandwidth 的需求非常高。回顧 Nvidia server,它其實有兩個分離的 networking domain:一個是 scale-up NVLink domain,另一個是 scale-out domain,後者通常使用 InfiniBand 或 Ethernet。

Atomic Claim 195/574 · 2026-02-16_inferencex-v2-nvidia-blackwell-vs::IX2-0202

Claim: EP 需要 all-to-all communication,也就是每顆 GPU 都必須把 token 傳送給其他每一顆 GPU
Frame: ATTRIBUTE · Mode: ASSERTED · Mapping: COMPLETE
開啟逐條審核

Atomic Claim 196/574 · 2026-02-16_inferencex-v2-nvidia-blackwell-vs::IX2-0204

Claim: 需要注意,Nvidia 的伺服器有兩個分離的 networking domains,其中之一是 scale-up NVLink domain。
Frame: ATTRIBUTE · Mode: ASSERTED · Mapping: COMPLETE
開啟逐條審核

Atomic Claim 197/574 · 2026-02-16_inferencex-v2-nvidia-blackwell-vs::IX2-0205

Claim: 另一個是 Scale-out Domain,通常採用 InfiniBandEthernet 作為 networking protocol。
Frame: RELATION · Mode: ASSERTED · Mapping: COMPLETE
開啟逐條審核

NVLink domain (within the NVL72 rack): 72 GPUs connected via NVLink with 900 GB/s uni-directional bandwidth per GPU. This is roughly 7-10x the bandwidth of the InfiniBand/Ethernet based scale-out network.

NVLink domain(NVL72 rack 內):72 顆 GPU 透過 NVLink 相連,每顆 GPU 單向 bandwidth 為 900 GB/s。這大約是 InfiniBand/Ethernet scale-out network bandwidth 的 7–10 倍。

Atomic Claim 198/574 · 2026-02-16_inferencex-v2-nvidia-blackwell-vs::IX2-0206

Claim: 在 NVL72 機櫃內的 NVLink domain 中,72 顆 GPUs 透過 NVLink 互連,每顆 GPU 提供 900 GB/s uni-directional bandwidth。
Frame: RELATION · Mode: ASSERTED · Mapping: COMPLETE
開啟逐條審核

Atomic Claim 199/574 · 2026-02-16_inferencex-v2-nvidia-blackwell-vs::IX2-0207

Claim: 這大約是基於 InfiniBandEthernetscale-out network 頻寬的 7~10 倍。
Frame: COMPARISON · Mode: ASSERTED · Mapping: COMPLETE
開啟逐條審核

InfiniBand/RoCEv2 Ethernet (outside of the NVL72 rack): Typically 400-800 Gbit/s per GPU uni-directional (50-100 GB/s). Note that all our testing for Nvidia was conducted on InfiniBand based clusters.

InfiniBand/RoCEv2 Ethernet(NVL72 rack 外):通常每顆 GPU 單向 400–800 Gbit/s,也就是 50–100 GB/s。要注意,我們所有 Nvidia 測試都是在 InfiniBand cluster 上完成。

Atomic Claim 200/574 · 2026-02-16_inferencex-v2-nvidia-blackwell-vs::IX2-0208

Claim: 在 NVL72 機櫃外,InfiniBandRoCEv2 Ethernet 通常每顆 GPU 提供 400~800 Gbit/s uni-directional,也就是約 50~100 GB/s
Frame: NARY_RELATION · Mode: ASSERTED · Mapping: PARTIAL
開啟逐條審核

Atomic Claim 201/574 · 2026-02-16_inferencex-v2-nvidia-blackwell-vs::IX2-0209

Claim: InferenceX 對 Nvidia 的所有測試都在以 InfiniBand 為基礎的 clusters 上進行。
Frame: ATTRIBUTE · Mode: ASSERTED · Mapping: PARTIAL
開啟逐條審核

TP shards every layer’s weight matrices across GPUs. This means that every single token at every single layer requires up to two all-reduce communications (one after the column-parallel GEMM, one after the row-parallel GEMM). For EP, all-to-all is done only at MoE layers. Each GPU sends only the tokens routed to each expert. This means cheaper comms across all layers for EP vs TP.

TP 會把每一層的 weight matrix shard 到多顆 GPU,因此每個 token、每一層最多都要做兩次 all-reduce:一次在 column-parallel GEMM 後、一次在 row-parallel GEMM 後。EP 則只在 MoE layer 做 all-to-all,而且每顆 GPU 只傳送被 route 到各 expert 的 token。也就是說,若看所有 layer 的 communication cost,EP 比 TP 更便宜。

Atomic Claim 202/574 · 2026-02-16_inferencex-v2-nvidia-blackwell-vs::IX2-0210

Claim: TP 會把每一層的 weight matrices 分片到多顆 GPUs
Frame: ATTRIBUTE · Mode: ASSERTED · Mapping: COMPLETE
開啟逐條審核

Atomic Claim 203/574 · 2026-02-16_inferencex-v2-nvidia-blackwell-vs::IX2-0211

Claim: 這代表每一層的每個 token 最多需要進行兩次 all-reduce 通訊,一次在 column-parallel GEMM 後、一次在 row-parallel GEMM 後。
Frame: ATTRIBUTE · Mode: ASSERTED · Mapping: COMPLETE
開啟逐條審核

Atomic Claim 204/574 · 2026-02-16_inferencex-v2-nvidia-blackwell-vs::IX2-0212

Claim: EP 則只在 MoE layers 執行 all-to-all。
Frame: ATTRIBUTE · Mode: ASSERTED · Mapping: COMPLETE
開啟逐條審核

Atomic Claim 205/574 · 2026-02-16_inferencex-v2-nvidia-blackwell-vs::IX2-0213

Claim: 每顆 GPU 只需傳送被 routing 到各 expert 的 token。
Frame: ATTRIBUTE · Mode: ASSERTED · Mapping: COMPLETE
開啟逐條審核

Atomic Claim 206/574 · 2026-02-16_inferencex-v2-nvidia-blackwell-vs::IX2-0214

Claim: 因此相較 TP,EP 在所有 layers 上的通訊成本較低。
Frame: COMPARISON · Mode: ASSERTED · Mapping: PARTIAL
開啟逐條審核

Because EP’s all-to-all communication bandwidth requirements scale with the number of participants, staying within the high-bandwidth NVLink domain before having to cross the slower IB/Eth fabric is better. With NVL72, EP across 72 GPUs is possible without ever leaving NVLink, whereas previous generations (with only 8-GPU NVLink domains) could only do EP across 8 GPUs at NVLink speed before hitting the slower IB/Eth networks.

EP 的 all-to-all bandwidth requirement 會隨 participant 數量增加,因此在不得不跨到較慢 IB/Ethernet fabric 之前,盡量留在高 bandwidth NVLink domain 內會比較好。NVL72 可以讓 72 顆 GPU 全部在 NVLink 內做 EP;前幾代只有 8-GPU NVLink domain,因此 EP 最多只能在 8 顆 GPU 上維持 NVLink speed,再往外就得進較慢的 IB/Ethernet network。

Atomic Claim 207/574 · 2026-02-16_inferencex-v2-nvidia-blackwell-vs::IX2-0215

Claim: 由於 EP 的 all-to-all communication 頻寬需求會隨參與者數量增加,因此最好盡可能留在高頻寬 NVLink domain 內,避免跨到較慢的 IB/Eth fabric。
Frame: COMPARISON · Mode: ASSERTED · Mapping: PARTIAL
開啟逐條審核

Atomic Claim 208/574 · 2026-02-16_inferencex-v2-nvidia-blackwell-vs::IX2-0216

Claim: 在 NVL72 上,可以讓 72 顆 GPUs 執行 EP,而完全不離開 NVLink domain。
Frame: ATTRIBUTE · Mode: ASSERTED · Mapping: COMPLETE
開啟逐條審核

Atomic Claim 209/574 · 2026-02-16_inferencex-v2-nvidia-blackwell-vs::IX2-0217

Claim: 上一代只有 8-GPU NVLink domains,因此 EP 最多只能在 8 顆 GPUs 內以 NVLink 速度運作,超過後就必須進入較慢的 IB/Eth networks。
Frame: COMPARISON · Mode: ASSERTED · Mapping: COMPLETE
開啟逐條審核

SemiAnalysis InferenceX is free open source software and reader-supported. To receive new posts and support our work, consider becoming a free or paid subscriber.

SemiAnalysis InferenceX 是免費 open-source software,也由讀者支持。如果想收到新文章並支持我們,可以考慮成為免費或付費訂閱者。

Subscribed

Wide EP also has a major advantage in weight loading efficiency. For a model like DeepSeek R1, decode is memory-bandwidth-bound: the bottleneck is how fast GPUs can load weights from HBM. With wide EP (e.g., DEP32), 32 GPUs collectively hold and load the 670B weights once, each loading only its shard (~21B). The total HBM bandwidth of all 32 chips is applied to loading a single copy of the model. By contrast, with narrower EP and more DP replicas (e.g., 5xDEP8), each of the 5 replicas needs its own full copy of the 670B weights, that’s 5×670B = 3.35T of redundant weight loading across the system. EP amortizes weights across chips; DP replicates them. This is why wider EP, enabled by high-bandwidth interconnects like NVLink, delivers significantly better throughput per GPU.

Wide EP 在 weight-loading efficiency 上也有很大優勢。像 DeepSeek R1 這類 model,decode 是 memory-bandwidth-bound,瓶頸在 GPU 能多快從 HBM 載入 weight。使用 wide EP,例如 DEP32,32 顆 GPU 共同持有並只載入一份 670B weight,每顆只負責自己的 shard(約 21B);32 顆 chip 的總 HBM bandwidth 都用來載入這一份 model。相反地,較窄 EP 搭配更多 DP replica,例如 5xDEP8,五個 replica 都需要自己的一整份 670B weight,也就是全 system 重複載入 5×670B = 3.35T。EP 把 weight cost amortize 到多顆 chip;DP 則直接複製 weight。這也是為什麼 NVLink 這種 high-bandwidth interconnect 所支援的 wider EP,可以顯著提高 throughput per GPU。

Atomic Claim 210/574 · 2026-02-16_inferencex-v2-nvidia-blackwell-vs::IX2-0218

Claim: Wide EP 在模型權重載入效率上也有重大優勢。
Frame: ATTRIBUTE · Mode: ASSERTED · Mapping: COMPLETE
開啟逐條審核

Atomic Claim 211/574 · 2026-02-16_inferencex-v2-nvidia-blackwell-vs::IX2-0219

Claim:DeepSeek R1 這類模型而言,decode 受到記憶體頻寬限制,瓶頸是 GPUsHBM 載入 weights 的速度。
Frame: ATTRIBUTE · Mode: ASSERTED · Mapping: PARTIAL
開啟逐條審核

Atomic Claim 212/574 · 2026-02-16_inferencex-v2-nvidia-blackwell-vs::IX2-0220

Claim: 使用 wide EP(例如 DEP32)時,32 顆 GPUs 共同持有並只載入一次 670B weights,每顆只需載入自己的 shard,約 21B。
Frame: ATTRIBUTE · Mode: ASSERTED · Mapping: COMPLETE
開啟逐條審核

Atomic Claim 213/574 · 2026-02-16_inferencex-v2-nvidia-blackwell-vs::IX2-0221

Claim: 32 顆晶片的總 HBM bandwidth 可以一起用來載入單一份模型。
Frame: ATTRIBUTE · Mode: ASSERTED · Mapping: COMPLETE
開啟逐條審核

Atomic Claim 214/574 · 2026-02-16_inferencex-v2-nvidia-blackwell-vs::IX2-0222

Claim: 相較之下,若使用較窄的 EP 與更多 DP replicas(例如 5×DEP8),5 個 replicas 每個都需要完整 670B weights,因此整個系統會重複載入 5×670B=3.35T 的權重。
Frame: ATTRIBUTE · Mode: HYPOTHETICAL · Mapping: PARTIAL
開啟逐條審核

Atomic Claim 215/574 · 2026-02-16_inferencex-v2-nvidia-blackwell-vs::IX2-0223

Claim: EP 可以把 weights 成本分攤到多顆晶片上。
Frame: ATTRIBUTE · Mode: ASSERTED · Mapping: COMPLETE
開啟逐條審核

Atomic Claim 216/574 · 2026-02-16_inferencex-v2-nvidia-blackwell-vs::IX2-0225

Claim: 因此高頻寬 NVLink 所支援的更寬 EP,能顯著提高每 GPU throughput。
Frame: RELATION · Mode: ASSERTED · Mapping: COMPLETE
開啟逐條審核

image

Source: SemiAnalysis InferenceX

Generally, TP is preferred at lower concurrencies due to load balancing. At small batch sizes, EP suffers from uneven token-to-expert routing, leaving some GPUs underutilized while others are overloaded. TP avoids this since each GPU holds a slice of every expert and always gets an equal share of work. At lower concurrency, the cost of this load imbalance outweighs TP’s additional communication overhead.

一般而言,lower concurrency 時 TP 會更受偏好,主要因為 load balancing。Small batch 下,EP 的 token-to-expert routing 容易不平均,造成有些 GPU 閒置、有些 GPU overload。TP 不會遇到同樣問題,因為每顆 GPU 都持有每個 expert 的一部分,工作量天然平均。Lower concurrency 時,這個 load imbalance 的成本往往大於 TP 額外 communication overhead。

Atomic Claim 217/574 · 2026-02-16_inferencex-v2-nvidia-blackwell-vs::IX2-0226

Claim: 一般而言,在較低 concurrency 下,由於 load balancing 考量,TP 會較受偏好。
Frame: COMPARISON · Mode: ASSERTED · Mapping: PARTIAL
開啟逐條審核

Atomic Claim 218/574 · 2026-02-16_inferencex-v2-nvidia-blackwell-vs::IX2-0227

Claim: 在小 batch size 下,EP 的 token-to-expert routing 容易不均,使部分 GPUs 利用率不足、其他 GPU 則過載。
Frame: ATTRIBUTE · Mode: ASSERTED · Mapping: COMPLETE
開啟逐條審核

Atomic Claim 219/574 · 2026-02-16_inferencex-v2-nvidia-blackwell-vs::IX2-0228

Claim: TP 可避免這個問題,因為每顆 GPU 都持有每個 expert 的一部分,因此工作量較平均。
Frame: ATTRIBUTE · Mode: ASSERTED · Mapping: COMPLETE
開啟逐條審核

Atomic Claim 220/574 · 2026-02-16_inferencex-v2-nvidia-blackwell-vs::IX2-0229

Claim: 在較低 concurrency 下,EP 的 load imbalance 成本會大於 TP 額外通訊開銷。
Frame: COMPARISON · Mode: ASSERTED · Mapping: PARTIAL
開啟逐條審核

At higher concurrencies, this tradeoff changes. Expert activation becomes more evenly distributed across larger batch sizes, and EP’s communication and weight-loading advantages dominate over TP’s expensive per-layer all-reduce. In the middle of the curve, hybrid TP+EP configurations balance both concerns using small TP groups within each expert for load balancing while EP is used across the wider set of GPUs to amortize weights and reduce communication.

到了 higher concurrency,trade-off 就會反轉。Batch 越大,expert activation 分布越平均,EP 在 communication 與 weight loading 的優勢開始超過 TP 每層昂貴 all-reduce 的代價。在 curve 中間區域,hybrid TP+EP configuration 可以兼顧兩者:每個 expert 內用小型 TP group 做 load balancing,再把 EP 擴到更大 GPU 集合,分攤 weight、降低 communication。

Atomic Claim 221/574 · 2026-02-16_inferencex-v2-nvidia-blackwell-vs::IX2-0231

Claim: 隨 batch size 變大,expert activation 會分布得更平均,此時 EP 在通訊與 weight loading 上的優勢會超越 TP 每層昂貴的 all-reduce 成本。
Frame: ATTRIBUTE · Mode: ASSERTED · Mapping: COMPLETE
開啟逐條審核

Atomic Claim 222/574 · 2026-02-16_inferencex-v2-nvidia-blackwell-vs::IX2-0232

Claim: 在曲線中段,hybrid TP+EP configuration 會兼顧兩者:在每個 expert 內使用小型 TP group 做 load balancing,同時在更大的 GPUs 集合間使用 EP,以分攤 weights 並降低通訊成本。
Frame: ATTRIBUTE · Mode: ASSERTED · Mapping: COMPLETE
開啟逐條審核

For higher interactivity levels (low batch size), large scale-up world sizes tend not to deliver stronger performance. B300 disagg over IB has the same performance as GB300 with NVL72, since the workload is latency-bound, not bandwidth-bound. The massive NVLink bandwidth advantage of NVL72 doesn’t matter because not even the much slower IB link is saturated by the tiny batches of tokens in flight.

在 higher interactivity、也就是 low batch size 區域,large scale-up world size 通常不會帶來更強 performance。B300 透過 IB 做 disagg,performance 和使用 NVL72 的 GB300 幾乎相同,因為這時 workload 是 latency-bound,不是 bandwidth-bound。NVL72 巨大的 NVLink bandwidth 優勢用不上,原因是 flight 中 token batch 太小,連慢很多的 IB link 都還沒吃滿。

Atomic Claim 223/574 · 2026-02-16_inferencex-v2-nvidia-blackwell-vs::IX2-0233

Claim: 在較高 interactivity、也就是低 batch size 情況下,大型 scale-up world size 通常不會帶來更強效能。
Frame: COMPARISON · Mode: ASSERTED · Mapping: PARTIAL
開啟逐條審核

Atomic Claim 224/574 · 2026-02-16_inferencex-v2-nvidia-blackwell-vs::IX2-0234

Claim: B300 透過 IB 執行 disagg 的效能與採 NVL72 的 GB300 相同,因為 workload 受 latency 限制,而不是 bandwidth 限制。
Frame: COMPARISON · Mode: ASSERTED · Mapping: COMPLETE
開啟逐條審核

Atomic Claim 225/574 · 2026-02-16_inferencex-v2-nvidia-blackwell-vs::IX2-0235

Claim: NVL72 巨大的 NVLink bandwidth 優勢在此無法發揮,因為流動中的 token batch 太小,連較慢的 IB link 都沒有被跑滿。
Frame: COMPARISON · Mode: ASSERTED · Mapping: PARTIAL
開啟逐條審核

Prefill/decode disaggregation also plays a role. Prefill is compute-heavy and bursty; decode is memory-bandwidth-bound and steady-state. When they share the same GPUs, they interfere with each other, causing latency jitter and wasted capacity. Separating them onto dedicated GPU pools lets each run a workload matched to its characteristics, improving effective utilization. This is why disaggregated B200 configs outperform single-node B200 in the middle of the throughput-interactivity curve. PD separation combined with wider EP across more GPUs over IB amortizes weights more efficiently than cramming both phases onto a single 8-GPU node.

Prefill/decode disaggregation 也很重要。Prefill compute-heavy 而且具有 burst 特性;decode 則 memory-bandwidth-bound、屬 steady-state。兩者共用同一批 GPU 時會互相干擾,帶來 latency jitter 與 capacity waste。把它們拆到 dedicated GPU pool,則能讓兩邊各自跑最符合自身特性的 workload,提高 effective utilization。這就是為什麼 throughput-interactivity curve 中段,disaggregated B200 configuration 會優於 single-node B200。PD separation 再搭配透過 IB、跨更多 GPU 的 wider EP,weight amortization efficiency 也比把兩個 phase 全塞進單一 8-GPU node 更好。

Atomic Claim 226/574 · 2026-02-16_inferencex-v2-nvidia-blackwell-vs::IX2-0236

Claim: Prefilldecode disaggregation 也會影響結果。
Frame: ATTRIBUTE · Mode: ASSERTED · Mapping: PARTIAL
開啟逐條審核

Atomic Claim 227/574 · 2026-02-16_inferencex-v2-nvidia-blackwell-vs::IX2-0237

Claim: Prefill 屬於運算密集且 bursty 的工作負載。
Frame: ATTRIBUTE · Mode: ASSERTED · Mapping: PARTIAL
開啟逐條審核

Atomic Claim 228/574 · 2026-02-16_inferencex-v2-nvidia-blackwell-vs::IX2-0238

Claim: decode 則是受記憶體頻寬限制、較穩定的 steady-state 工作負載。
Frame: ATTRIBUTE · Mode: ASSERTED · Mapping: COMPLETE
開啟逐條審核

Atomic Claim 229/574 · 2026-02-16_inferencex-v2-nvidia-blackwell-vs::IX2-0239

Claim: 當兩者共用同一批 GPUs 時,彼此會互相干擾,造成 latency jitter 與容量浪費。
Frame: ATTRIBUTE · Mode: ASSERTED · Mapping: COMPLETE
開啟逐條審核

Atomic Claim 230/574 · 2026-02-16_inferencex-v2-nvidia-blackwell-vs::IX2-0240

Claim: 將兩者分配到專用 GPU pools 後,每個 pool 都能執行更符合自身特性的 workload,進而提高有效利用率。
Frame: ATTRIBUTE · Mode: ASSERTED · Mapping: COMPLETE
開啟逐條審核

Atomic Claim 231/574 · 2026-02-16_inferencex-v2-nvidia-blackwell-vs::IX2-0241

Claim: 這也是為什麼 disaggregated B200 configurations 在 throughput-interactivity curve 中段會優於 single-node B200
Frame: COMPARISON · Mode: ASSERTED · Mapping: PARTIAL
開啟逐條審核

Atomic Claim 232/574 · 2026-02-16_inferencex-v2-nvidia-blackwell-vs::IX2-0242

Claim: PD separation 搭配更多 GPUs 上的 wider EP,即使跨 IB,也能比把 prefilldecode 都塞進單一 8-GPU node 更有效地分攤 weights。
Frame: ATTRIBUTE · Mode: ASSERTED · Mapping: PARTIAL
開啟逐條審核

Side Note: the 10x inference engineers at TogetherAI noticed an pattern for multi-turn traffic where the requirements of first turn prefill is much different from the following turns prefill’s and disaggregrated it leading to better TTFT performance.

補充:TogetherAI 的 10x inference engineer 發現 multi-turn traffic 有一個特性——first-turn prefill 的需求和後續 turn 的 prefill 很不一樣,因此把兩者進一步 disaggregate,能改善 TTFT performance ↗。

Atomic Claim 233/574 · 2026-02-16_inferencex-v2-nvidia-blackwell-vs::IX2-0243

Claim: 補充:TogetherAI 的 inference engineers 發現 multi-turn traffic 中,第一輪 prefill 與後續輪次 prefill 的需求差異很大,因此將其進一步 disaggregate 後,可改善 TTFT
Frame: ATTRIBUTE · Mode: ASSERTED · Mapping: COMPLETE
開啟逐條審核

image

Source: SemiAnalysis InferenceX

Jensen Under Promising and Overdelivering - Hopper vs Blackwell vs Rack Scale NVL72

At GTC 2024, Jensen was on stage promising up to 30x performance gains from H100 to GB200 NVL72, everyone thought it was classic marketing lookmaxxing and would not be achievable in real world. Many looked to come up with labels for this perceived use of a reality distortion field so they could crack more Jensen Math jokes. Indeed – we did point to the comparison of 30x performance difference between the worst case for H200 on FP8 to a reasonable case of the GB200 on FP4.

GTC 2024 時,Jensen 在台上宣稱從 H100 到 GB200 NVL72 最多可提升 30x performance,當時大家都覺得這是典型 marketing lookmaxxing,real world 根本做不到 ↗。不少人還替這種 perceived reality-distortion field 想各種名稱,好繼續開「Jensen Math」玩笑。當時我們也確實指出,所謂 30x 是拿 H200 FP8 的 worst case ↗,去比 GB200 FP4 一個合理 case。

Atomic Claim 234/574 · 2026-02-16_inferencex-v2-nvidia-blackwell-vs::IX2-0244

Claim: GTC 2024 時,Jensen 曾在台上宣稱從 H100 升級到 GB200 NVL72 最多可獲得 30 倍效能提升,當時許多人認為這只是行銷宣傳、現實中難以達成。
Frame: COMPARISON · Mode: ATTRIBUTED · Mapping: COMPLETE
開啟逐條審核

Atomic Claim 235/574 · 2026-02-16_inferencex-v2-nvidia-blackwell-vs::IX2-0246

Claim: SemiAnalysis 當時也曾指出,30 倍比較是拿 H200 FP8 的較差情境,對比 GB200 FP4 的合理情境。
Frame: COMPARISON · Mode: ASSERTED · Mapping: PARTIAL
開啟逐條審核

Atomic Claim 236/574 · 2026-02-16_inferencex-v2-nvidia-blackwell-vs::IX2-0245

Claim: 很多人甚至試圖為這種被認為是「reality distortion field」的宣傳方式取名,以便繼續開 Jensen Math 的玩笑。
Frame: CLAIM_ONLY · Mode: ASSERTED · Mapping: CLAIM_ONLY
開啟逐條審核

Nvidia Blackwell Perf TCO Analysis - B100 vs B200 vs GB200NVL72

Nvidia Blackwell Perf TCO Analysis - B100 vs B200 vs GB200NVL72

Dylan Patel and Daniel Nishball

Dylan Patel ↗、Daniel Nishball ↗

Atomic Claim 237/574 · 2026-02-16_inferencex-v2-nvidia-blackwell-vs::IX2-0247

Claim: 此處署名為 Dylan Patel 與 Daniel Nishball。
Frame: CLAIM_ONLY · Mode: ASSERTED · Mapping: CLAIM_ONLY
開啟逐條審核

·

2024年4月10日

Read full story

閱讀完整文章 ↗

image

Source: Nvidia GTC 2024

But it turns out the joke is on them. Fast forward almost two years later, and we can now see that it wasn’t marketing hype lookmaxing after all, and Jensen was actually under promising on Blackwell performance the whole time. From our testing, Blackwell is so good at large scale MoE inferencing compared to even a strong H100 disagg+wideEP FP8 baseline that it, at 116 toks/s/user, delivers up to 98x better perf on GB200 NVL72 FP4 and up to 100x better perf on GB300 NVL72 FP4! Maybe the new Jensen Math rule is that he delivers double whatever he promises in terms of token throughput. The more you spend, the more you save indeed!

結果最後被打臉的是懷疑的人。快轉將近兩年,我們現在可以看到那根本不只是 marketing hype lookmaxxing;Jensen 在 Blackwell performance 上反而一直是 under-promise。根據我們測試,Blackwell 做 large-scale MoE inference 相較甚至很強的 H100 disagg+wideEP FP8 baseline 都好太多:在 116 tok/s/user 下,GB200 NVL72 FP4 最多可提升 98x,GB300 NVL72 FP4 最多甚至達 100x。也許新版 Jensen Math 規則是:token throughput 最後會交付承諾值的兩倍。The more you spend, the more you save,還真的成立了。

Atomic Claim 238/574 · 2026-02-16_inferencex-v2-nvidia-blackwell-vs::IX2-0248

Claim: 但後來看來,真正被打臉的是當時嘲笑這些數字的人。
Frame: CLAIM_ONLY · Mode: ASSERTED · Mapping: CLAIM_ONLY
開啟逐條審核

Atomic Claim 239/574 · 2026-02-16_inferencex-v2-nvidia-blackwell-vs::IX2-0251

Claim: Jensen 實際上一直低估了 Blackwell 的效能。
Frame: ATTRIBUTE · Mode: ASSERTED · Mapping: PARTIAL
開啟逐條審核

Atomic Claim 240/574 · 2026-02-16_inferencex-v2-nvidia-blackwell-vs::IX2-0252

Claim: 根據 InferenceX 測試,Blackwell 在大規模 MoE 推論上相較即使是很強的 H100 disagg+wideEP FP8 baseline 仍有巨大優勢;在 116 toks/s/user 下,GB200 NVL72 FP4 最高可達 98 倍,GB300 NVL72 FP4 最高可達 100 倍效能。
Frame: COMPARISON · Mode: ASSERTED · Mapping: PARTIAL
開啟逐條審核

Atomic Claim 241/574 · 2026-02-16_inferencex-v2-nvidia-blackwell-vs::IX2-0253

Claim: SemiAnalysis 戲稱新的 Jensen Math 規則可能是:token throughput 最後會做到承諾值的兩倍。
Frame: CLAIM_ONLY · Mode: ASSERTED · Mapping: CLAIM_ONLY
開啟逐條審核

image

Source: SemiAnalysis InferenceX

Even when factoring in the increased total cost of ownership of Blackwell and Blackwell Ultra, we see a 9.7x(40 tok/s/user) up to 65x(116 tok/s/user) improvement in tokens per dollar compared to Hopper. You can explore Hopper vs Blackwell performance in detail on our free website . Blackwell performance is so good compared to Hopper that we needed to an log scale to our dashboard in order to visualize it.

即使把 Blackwell、Blackwell Ultra 更高的 total cost of ownership 算進去,相較 Hopper,tokens per dollar 仍可改善約 9.7x(40 tok/s/user)到 65x(116 tok/s/user)。Hopper vs Blackwell 的詳細 performance 可在我們免費網站自行探索 ↗。Blackwell 相較 Hopper 的差距大到,我們甚至必須替 dashboard 加上 log scale 才能正常顯示。

Atomic Claim 242/574 · 2026-02-16_inferencex-v2-nvidia-blackwell-vs::IX2-0255

Claim: 即使把 BlackwellBlackwell Ultra 較高的 TCO 納入考量,相較 Hopper,tokens per dollar 仍改善約 9.7 倍(40 tok/s/user)到 65 倍(116 tok/s/user)。
Frame: COMPARISON · Mode: ASSERTED · Mapping: COMPLETE
開啟逐條審核

Atomic Claim 243/574 · 2026-02-16_inferencex-v2-nvidia-blackwell-vs::IX2-0256

Claim: 使用者可在 InferenceX 免費網站上深入探索 HopperBlackwell 的效能差異。
Frame: RELATION · Mode: ASSERTED · Mapping: PARTIAL
開啟逐條審核

Atomic Claim 244/574 · 2026-02-16_inferencex-v2-nvidia-blackwell-vs::IX2-0257

Claim: Blackwell 相較 Hopper 的效能差距大到 InferenceX dashboard 必須加入 log scale 才方便視覺化。
Frame: COMPARISON · Mode: ASSERTED · Mapping: PARTIAL
開啟逐條審核

image

Source: SemiAnalysis InferenceX

As mentioned earlier in the article, B300 servers only connect at most 8 GPUs using the 900GByte/s/GPU NVLink scale-up network whereas GB300 NVL72 servers connect 72 GPUs using the NVlink scale-up network. So when we need more than 8 GPUs (but less than 72 GPUs) for the inference setup, we need to bring in multiple nodes of B300 servers to form our inference system which means communications falls back to the lower InfiniBand XDR scale-out network featuring 800Gbit/s (uni-di) per GPU of bandwidth. Compare this to a rack scale GB300 NVL72 which connects 72 GPUs over NVLink delivering 900GByte/s (uni-di) per GPU of bandwidth and we can see that the rack-scale server allows the GPUs in the inference setup to talk to each other with over 9x higher bandwidth compared to the case of the multiple nodes of B300 servers.

如前文所述,B300 server 的 900 GByte/s/GPU NVLink scale-up network 最多只連 8 顆 GPU;GB300 NVL72 則能用 NVLink 連 72 顆 GPU。因此 inference setup 若需要超過 8 顆、但少於 72 顆 GPU,B300 必須組多個 node,communication 就會退回 bandwidth 較低的 InfiniBand XDR scale-out network,每顆 GPU 單向約 800 Gbit/s。相比之下,rack-scale GB300 NVL72 的 72 顆 GPU 都透過 NVLink 相連,每顆 GPU 單向 900 GByte/s。也就是說,在 inference setup 內 GPU 彼此 communication 的 bandwidth,rack-scale server 比 multi-node B300 高超過 9x。

Atomic Claim 245/574 · 2026-02-16_inferencex-v2-nvidia-blackwell-vs::IX2-0258

Claim: 如前文所述,B300 server 最多只能用 900GByte/s/GPUNVLink scale-up network 連接 8 顆 GPUs,而 GB300 NVL72 server 可在同一 NVlink scale-up network 內連接 72 顆 GPUs
Frame: NARY_RELATION · Mode: ASSERTED · Mapping: PARTIAL
開啟逐條審核

Atomic Claim 246/574 · 2026-02-16_inferencex-v2-nvidia-blackwell-vs::IX2-0259

Claim: 因此當推論配置需要超過 8 顆 GPUs、但少於 72 顆 GPUs 時,B300 必須使用多個 nodes 組成系統,通訊就會退回頻寬較低的 InfiniBand XDR scale-out network,每顆 GPU 僅有 800Gbit/s uni-di。
Frame: COMPARISON · Mode: ASSERTED · Mapping: PARTIAL
開啟逐條審核

Atomic Claim 247/574 · 2026-02-16_inferencex-v2-nvidia-blackwell-vs::IX2-0260

Claim: 相較之下,rack-scale GB300 NVL72 可透過 NVLink 連接 72 顆 GPUs,每顆 GPU 提供 900GByte/s uni-di,因此推論系統中的 GPUs 彼此通訊頻寬比多節點 B300 高超過 9 倍。
Frame: COMPARISON · Mode: ASSERTED · Mapping: PARTIAL
開啟逐條審核

SemiAnalysis is free open source software and reader-supported. To receive new posts and support our work, consider becoming a free or paid subscriber.

SemiAnalysis 是免費 open-source software,也由讀者支持。如果想收到新文章並支持我們,可以考慮成為免費或付費訂閱者。

Atomic Claim 248/574 · 2026-02-16_inferencex-v2-nvidia-blackwell-vs::IX2-0261

Claim: SemiAnalysis 的相關軟體為免費開源,並由讀者支持。
Frame: CLAIM_ONLY · Mode: ASSERTED · Mapping: CLAIM_ONLY
開啟逐條審核

Subscribed

image

Source: SemiAnalysis InferenceX

Admittedly the GB300 NVL72 has a higher all-in cost per GPU, but this only reduces the bandwidth per TCO advantage to being 8x faster. The bandwidth advantage of the rack-scale architecture directly drives a much lower cost per token. Google TPU, AWS Trainium and Nvidia are the only AI chips to have rack scale system designs deployed today. Engineering samples and low volume production of AMD’s first rack scale MI455X UALoE72 system will be in H2 2026 while due to manufacturing delays, the mass production ramp and first production tokens will only be generated on an MI455X UALoE72 by Q2 2027.

確實,GB300 NVL72 的 all-in cost per GPU 比較高,但把成本算進去後,bandwidth per TCO 優勢仍約 8x。Rack-scale architecture 的 bandwidth 優勢,會直接轉化成更低 cost per token。目前真正已有 rack-scale system design 部署的 AI chip,只有 Google TPU、AWS Trainium 與 Nvidia。AMD 第一套 rack-scale MI455X UALoE72,engineering sample 與 low-volume production 預計在 H2 2026;但因 manufacturing delay,mass-production ramp 與第一批 production token 要到 Q2 2027 才會真正產生。

Atomic Claim 249/574 · 2026-02-16_inferencex-v2-nvidia-blackwell-vs::IX2-0262

Claim: 雖然 GB300 NVL72 每顆 GPU 的 all-in cost 較高,但折算成 bandwidth per TCO 後,仍有約 8 倍速度優勢。
Frame: COMPARISON · Mode: ASSERTED · Mapping: COMPLETE
開啟逐條審核

Atomic Claim 250/574 · 2026-02-16_inferencex-v2-nvidia-blackwell-vs::IX2-0263

Claim: rack-scale architecture 的頻寬優勢會直接轉化成更低的每 token 成本。
Frame: ATTRIBUTE · Mode: ASSERTED · Mapping: PARTIAL
開啟逐條審核

Atomic Claim 251/574 · 2026-02-16_inferencex-v2-nvidia-blackwell-vs::IX2-0264

Claim: 目前已實際部署 rack-scale system designs 的 AI chips 只有 Google TPUAWS TrainiumNvidia
Frame: NARY_RELATION · Mode: ASSERTED · Mapping: PARTIAL
開啟逐條審核

Atomic Claim 252/574 · 2026-02-16_inferencex-v2-nvidia-blackwell-vs::IX2-0265

Claim: AMD MI455X UALoE72 的 engineering samples 與低量生產預計在 2026 年下半年開始。
Frame: ATTRIBUTE · Mode: ASSERTED · Mapping: PARTIAL
開啟逐條審核

Atomic Claim 253/574 · 2026-02-16_inferencex-v2-nvidia-blackwell-vs::IX2-0266

Claim: 由於製造延遲,MI455X UALoE72 mass production 預計要到 2027 年第二季。
Frame: ATTRIBUTE · Mode: ASSERTED · Mapping: COMPLETE
開啟逐條審核

Atomic Claim 254/574 · 2026-02-16_inferencex-v2-nvidia-blackwell-vs::IX2-0267

Claim: MI455X UALoE72 真正產出 production tokens 預計也要到 2027 年第二季。
Frame: ATTRIBUTE · Mode: ASSERTED · Mapping: COMPLETE
開啟逐條審核

image

Source: SemiAnalysis InferenceX

Blackwell vs Blackwell Ultra

On paper, the newly released Blackwell Ultra has the same memory bandwidth as Blackwell, the same FP8 performance and only 1.5x higher FP4 performance, but when measuring we actually see up to 1.5x better FP8 performance on the Blackwell Ultra, though we only see 1.1x better performance on FP4. This may be due to Blackwell Ultra being a newly released GPU, meaning software is not fully optimized yet.

規格表上,新推出的 Blackwell Ultra 和 Blackwell 有相同 memory bandwidth、相同 FP8 performance,FP4 performance 只有 1.5x;但實測反而看到 Blackwell Ultra 的 FP8 最多提升 1.5x,FP4 卻只有約 1.1x。可能原因是 Blackwell Ultra 才剛發布,software optimization 還沒成熟。

Atomic Claim 255/574 · 2026-02-16_inferencex-v2-nvidia-blackwell-vs::IX2-0270

Claim: 從紙面規格看,新推出的 Blackwell Ultra 的 FP4 效能只高約 1.5 倍。
Frame: COMPARISON · Mode: ASSERTED · Mapping: COMPLETE
開啟逐條審核

Atomic Claim 256/574 · 2026-02-16_inferencex-v2-nvidia-blackwell-vs::IX2-0272

Claim: 在紙面規格上,新發布的 Blackwell Ultra 有所提升,但實測 FP4 效能只高約 1.1 倍。
Frame: COMPARISON · Mode: ASSERTED · Mapping: COMPLETE
開啟逐條審核

Atomic Claim 257/574 · 2026-02-16_inferencex-v2-nvidia-blackwell-vs::IX2-0271

Claim: 但實際測量時,Blackwell Ultra 的 FP8 效能最高可比 Blackwell 高 1.5 倍。
Frame: COMPARISON · Mode: ASSERTED · Mapping: COMPLETE
開啟逐條審核

Atomic Claim 258/574 · 2026-02-16_inferencex-v2-nvidia-blackwell-vs::IX2-0269

Claim: 從紙面規格看,新推出的 Blackwell Ultra 與 Blackwell 有相同 FP8 效能。
Frame: COMPARISON · Mode: ASSERTED · Mapping: COMPLETE
開啟逐條審核

Atomic Claim 259/574 · 2026-02-16_inferencex-v2-nvidia-blackwell-vs::IX2-0268

Claim: 從紙面規格看,新推出的 Blackwell Ultra 與 Blackwell 有相同記憶體頻寬。
Frame: COMPARISON · Mode: ASSERTED · Mapping: PARTIAL
開啟逐條審核

Atomic Claim 260/574 · 2026-02-16_inferencex-v2-nvidia-blackwell-vs::IX2-0273

Claim: 這可能因為 Blackwell Ultra 是剛推出的 GPU,software 尚未完全最佳化。
Frame: RELATION · Mode: HYPOTHETICAL · Mapping: PARTIAL
開啟逐條審核

image

image

Source: SemiAnalysis InferenceX

MI355X vs MI325X vs MI300X

On AMD SKUs, we see up to 10x better performance on the MI355X vs the MI300X. AMD has only gotten DeepSeek SGLang Disaggregated Inferencing to work on the MI355X so far AMD has not submitted MI300X or MI325X disaggregated inferencing results, potentially due to software issues on older SKUs that are still being solved.

AMD SKU 方面,MI355X 相較 MI300X 最多可看到約 10x performance improvement。目前 AMD 只把 DeepSeek SGLang Disaggregated Inference 在 MI355X 上跑通,MI300X、MI325X 的 disaggregated result 還沒有提交,可能是舊 SKU software issue 仍在處理。

Atomic Claim 261/574 · 2026-02-16_inferencex-v2-nvidia-blackwell-vs::IX2-0274

Claim:AMD SKUs 中,MI355X 相較 MI300X 最高可有約 10 倍效能提升。
Frame: COMPARISON · Mode: ASSERTED · Mapping: COMPLETE
開啟逐條審核

Atomic Claim 262/574 · 2026-02-16_inferencex-v2-nvidia-blackwell-vs::IX2-0275

Claim: 目前 AMD 只讓 DeepSeek SGLang Disaggregated Inferencing 在 MI355X 上正常運作;AMD 尚未提交 MI300XMI325X 的 disaggregated inference 結果,可能是較舊 SKUs 仍有 software issues 尚待解決。
Frame: ATTRIBUTE · Mode: ASSERTED · Mapping: COMPLETE
開啟逐條審核

image

Source: SemiAnalysis InferenceX

image

image

Source: SemiAnalysis InferenceX

Turning to cost, for DeepSeekR1 on FP8, at an interactivity of 24 tok/s/user, the MI355X delivers inferences a cost that is slightly less than 3x cheaper than for the MI325X. The throughput of each GPU is slightly less than 4 times that of MI325X.

從 cost 看 DeepSeek R1 FP8,在 24 tok/s/user interactivity 下,MI355X inference cost 比 MI325X 便宜接近 3x;每顆 GPU throughput 則略低於 MI325X 的 4 倍。

Atomic Claim 263/574 · 2026-02-16_inferencex-v2-nvidia-blackwell-vs::IX2-0276

Claim: 就成本而言,在 DeepSeekR1 FP8、24 tok/s/user interactivity 下,MI355X 的推論成本略低於 MI325X 的三分之一。
Frame: COMPARISON · Mode: ASSERTED · Mapping: PARTIAL
開啟逐條審核

Atomic Claim 264/574 · 2026-02-16_inferencex-v2-nvidia-blackwell-vs::IX2-0277

Claim: 每顆 GPU 的 throughput 略低於 MI325X 的 4 倍。
Frame: COMPARISON · Mode: ASSERTED · Mapping: COMPLETE
開啟逐條審核

image

Source: SemiAnalysis InferenceX

AMD Composability Issue on FP4, Distributed Inferencing and Wide Expert Parallelism

While AMD performs somewhat decently on single node FP4 and performs competitively to B200 SGLang on FP8 distributed inferencing, the issue with the current AMD open source inferencing stack is that, while individual inference optimizations perform well, real customers deploy with multiple optimizations composed together. Top tier AI labs are all using FP4 **with **disaggregated inferencing with wide expert parallelism all enabled at the same time, and this is where the issue occurs.

AMD 在 single-node FP4 表現還算可以,FP8 distributed inference 也能和 B200 SGLang 競爭;但目前 AMD open-source inference stack 的問題是:單一 optimization 各自看都不差,可是 real customer 會把多種 optimization 疊在一起用。Top-tier AI lab 現在全部都是 FP4、disaggregated inference、wide expert parallelism 同時啟用,而 AMD 真正的問題正是在這個組合情境出現。

Atomic Claim 265/574 · 2026-02-16_inferencex-v2-nvidia-blackwell-vs::IX2-0278

Claim: 雖然 AMD 在 single-node FP4 表現尚可,FP8 distributed inferencing 搭配 SGLang 時也能與 B200 競爭,但 AMD 目前 open-source inference stack 的核心問題仍存在。
Frame: ATTRIBUTE · Mode: ASSERTED · Mapping: COMPLETE
開啟逐條審核

Atomic Claim 266/574 · 2026-02-16_inferencex-v2-nvidia-blackwell-vs::IX2-0280

Claim: 頂級 AI labs 已經同時啟用 FP4、disaggregated inferencing 與 wide expert parallelism
Frame: ATTRIBUTE · Mode: ASSERTED · Mapping: PARTIAL
開啟逐條審核

Atomic Claim 267/574 · 2026-02-16_inferencex-v2-nvidia-blackwell-vs::IX2-0281

Claim: 關於 EP:問題正是在這些最佳化同時組合時出現。
Frame: ATTRIBUTE · Mode: ASSERTED · Mapping: COMPLETE
開啟逐條審核

SemiAnalysis is free open source software and reader-supported. To receive new posts and support our work, consider becoming a free or paid subscriber.

SemiAnalysis 是免費 open-source software,也由讀者支持。如果想收到新文章並支持我們,可以考慮成為免費或付費訂閱者。

Atomic Claim 268/574 · 2026-02-16_inferencex-v2-nvidia-blackwell-vs::IX2-0282

Claim: SemiAnalysis 的相關軟體為免費開源,並由讀者支持。
Frame: CLAIM_ONLY · Mode: ASSERTED · Mapping: CLAIM_ONLY
開啟逐條審核

Subscribed

AMD software is still not meeting the mark, and the theoretical speed of light modelling at SemiAnalysis and at AMD show that for FP4, disaggregated inferencing with wide expert parallelism should perform better than inference on a single node of MI355X. Unfortunately, Software continues to be a massive bottleneck for AMD GPUs. AMD management needs to continue to sharpen resource allocation of their engineering talent, for instance, re-allocate their engineering resources away from pet single node projects that nobody uses like ATOM towards fixing the aforementioned issues with composability of inference optimizations between disaggregated inferencing, wide expert parallelism and FP4. The current subpar software is due to lack of focus and incorrect prioritization of where the industry already is at. All top tier labs are already using disaggregated inferencing and wide expert parallelism; AMD needs to stop focusing on single node and heavily invest focus into multi node inferencing for open source solutions.

AMD software 目前仍未達標。SemiAnalysis 與 AMD 自己的 theoretical speed-of-light modeling 都顯示,FP4 + disaggregated inference + wide expert parallelism 理論上應該優於 MI355X single-node inference,但 software 依然是 AMD GPU 的巨大 bottleneck。AMD management 應持續改善 engineering talent 的 resource allocation,例如把資源從幾乎沒有人真正使用的 single-node pet project(像 ATOM)移開,投入前面提到的 disaggregated inference、wide expert parallelism、FP4 之間的 composability。現在 software 偏弱,本質是 focus 不夠、priority 和產業實際進度錯位。所有 top-tier lab 都已經在用 disaggregated inference + wide expert parallelism;AMD 應停止過度聚焦 single-node,大幅投資 open-source multi-node inference。

Atomic Claim 269/574 · 2026-02-16_inferencex-v2-nvidia-blackwell-vs::IX2-0283

Claim: SemiAnalysis 認為 AMD software 目前仍未達到應有水準。
Frame: ATTRIBUTE · Mode: ASSERTED · Mapping: COMPLETE
開啟逐條審核

Atomic Claim 270/574 · 2026-02-16_inferencex-v2-nvidia-blackwell-vs::IX2-0284

Claim: SemiAnalysis 與 AMD 的 theoretical speed-of-light modeling 都顯示,在 FP4 下,搭配 wide expert parallelism 的 disaggregated inference 理論上應該優於單一 MI355X node。
Frame: COMPARISON · Mode: EXPECTED · Mapping: COMPLETE
開啟逐條審核

Atomic Claim 271/574 · 2026-02-16_inferencex-v2-nvidia-blackwell-vs::IX2-0285

Claim: 但 software 仍然是 AMD GPUs 的重大瓶頸。
Frame: ATTRIBUTE · Mode: ASSERTED · Mapping: COMPLETE
開啟逐條審核

Atomic Claim 272/574 · 2026-02-16_inferencex-v2-nvidia-blackwell-vs::IX2-0286

Claim: SemiAnalysis 認為 AMD management 應進一步改善工程人才配置,例如把資源從幾乎沒有客戶使用的 single-node pet projects(如 ATOM)轉向解決 disaggregated inference、wide expert parallelismFP4 之間的 composability 問題。
Frame: NARY_RELATION · Mode: ASSERTED · Mapping: COMPLETE
開啟逐條審核

Atomic Claim 273/574 · 2026-02-16_inferencex-v2-nvidia-blackwell-vs::IX2-0288

Claim: 所有頂級 labs 都已使用 disaggregated inferencing 與 wide expert parallelism
Frame: ATTRIBUTE · Mode: ASSERTED · Mapping: PARTIAL
開啟逐條審核

Atomic Claim 274/574 · 2026-02-16_inferencex-v2-nvidia-blackwell-vs::IX2-0289

Claim: AMD 應停止把重心放在 single-node,並大幅增加對 open-source multi-node inferencing 的投入。
Frame: ATTRIBUTE · Mode: ASSERTED · Mapping: PARTIAL
開啟逐條審核

AMD is more than six months behind on open source distributed inferencing and wide expert parallelism and FP4 composability as shown by Nvidia and SGLang team showing off their NVFP4 performance on DeepSeek six months ago .

AMD 在 open-source distributed inference、wide expert parallelism 與 FP4 composability 上,已落後超過六個月;Nvidia 與 SGLang team 六個月前就已展示 DeepSeek NVFP4 performance ↗。

Atomic Claim 275/574 · 2026-02-16_inferencex-v2-nvidia-blackwell-vs::IX2-0290

Claim: AMD 在 open-source distributed inferencing、wide expert parallelismFP4 composability 上落後超過六個月;NvidiaSGLang 團隊早在六個月前就已展示 DeepSeek 上的 NVFP4 效能。
Frame: ATTRIBUTE · Mode: ASSERTED · Mapping: COMPLETE
開啟逐條審核

image

Source: SemiAnalysis InferenceX

AMD ATOM Engine

AMD has launched a new inference engine called ATOM. Atom can deliver slightly better single node performance, but it is completely lacking on a lot of features that makes it unusable for real workloads. One such example is that it does not support NVMe or CPU KVCache offloading, tool parsing, wide expert parallelism, or disaggregated serving. This has led to zero customers using it in production. Unlike Nvidia’s TRTLLM which generates billions of tokens per hour globally at companies like TogetherAI, etc and does support tool parsing and other features , there are no token factories currently using ATOM due to the lack of the aforementioned features.

AMD 推出了一套新的 inference engine:ATOM。ATOM 的 single-node performance 可以稍微更好,但缺少大量 real workload 必需功能,導致實務上幾乎不可用。例如它不支援 NVMe/CPU KVCache offloading、tool parsing、wide expert parallelism、disaggregated serving,因此目前 production customer 是零。相比 Nvidia TRTLLM 已經在 TogetherAI 等公司每小時產生數十億 token,並支援 tool parsing 等功能 ↗,目前沒有任何 token factory 使用 ATOM,主因就是上述功能缺口。

Atomic Claim 276/574 · 2026-02-16_inferencex-v2-nvidia-blackwell-vs::IX2-0296

Claim: 相較之下,NvidiaTRTLLM 支援 tool parsing 等功能,並已在 TogetherAI 等公司全球每小時產生數十億 tokens;ATOM 因缺少前述功能,目前沒有 token factory 使用。
Frame: COMPARISON · Mode: ASSERTED · Mapping: PARTIAL
開啟逐條審核

Atomic Claim 277/574 · 2026-02-16_inferencex-v2-nvidia-blackwell-vs::IX2-0291

Claim: AMD 推出一套名為 ATOM 的新 inference engine
Frame: NARY_RELATION · Mode: ASSERTED · Mapping: COMPLETE
開啟逐條審核

Atomic Claim 278/574 · 2026-02-16_inferencex-v2-nvidia-blackwell-vs::IX2-0292

Claim: ATOM 在 single-node 上可以提供略好的效能。
Frame: ATTRIBUTE · Mode: ASSERTED · Mapping: PARTIAL
開啟逐條審核

Atomic Claim 279/574 · 2026-02-16_inferencex-v2-nvidia-blackwell-vs::IX2-0293

Claim:ATOM 缺少許多必要功能,因此無法用於真實 production workloads。
Frame: ATTRIBUTE · Mode: ASSERTED · Mapping: COMPLETE
開啟逐條審核

Atomic Claim 280/574 · 2026-02-16_inferencex-v2-nvidia-blackwell-vs::IX2-0294

Claim: 例如 ATOM 不支援 NVMeCPU KVCache offloading、tool parsing、wide expert parallelismdisaggregated serving
Frame: NARY_RELATION · Mode: ASSERTED · Mapping: COMPLETE
開啟逐條審核

Atomic Claim 281/574 · 2026-02-16_inferencex-v2-nvidia-blackwell-vs::IX2-0295

Claim: 因此目前沒有客戶在 production 使用 ATOM
Frame: ATTRIBUTE · Mode: ASSERTED · Mapping: COMPLETE
開啟逐條審核

Furthermore, maintainers of open-source inference engines like vLLM are disappointed in AMD due to a lack of engineering and GPU resources provided by AMD. For example, Simon Mo, lead vLLM maintainer, states in this GitHub RFC that there is still no working MI355X that he can add to vLLM CI, hence the poor user experience. There are currently zero Mi355X tests on vLLM, while NVIDIA’s B200 has many tests on vLLM. Similarly, there are still not enough MI300X CI machines on vLLM. Upstream vLLM needs at least 20 more MI300 machines, 20 more MI325 machines and 20 more MI355X machines to reach the same level of usability as CUDA.

此外,vLLM 等 open-source inference engine 的 maintainer 也對 AMD 不滿,原因是 AMD 提供的 engineering 與 GPU resource 太少。例如 vLLM lead maintainer Simon Mo 在 GitHub RFC 指出,他到現在仍拿不到一台可正常加入 vLLM CI 的 MI355X,user experience 因此受到影響。目前 vLLM 對 MI355X 的 test 是零,但 NVIDIA B200 已有大量 test;MI300X CI machine 也仍不夠。若要讓 upstream vLLM 的 usability 接近 CUDA,至少還需要再增加約 20 台 MI300、20 台 MI325、20 台 MI355X。

Atomic Claim 282/574 · 2026-02-16_inferencex-v2-nvidia-blackwell-vs::IX2-0297

Claim: 此外,vLLM 等 open-source inference engine maintainers 對 AMD 提供的工程與 GPU 資源不足感到失望,並將此歸因於 AMD
Frame: NARY_RELATION · Mode: ASSERTED · Mapping: PARTIAL
開啟逐條審核

Atomic Claim 283/574 · 2026-02-16_inferencex-v2-nvidia-blackwell-vs::IX2-0298

Claim: 例如 vLLM lead maintainer Simon Mo 在 GitHub RFC 中指出,目前仍沒有一台可正常工作的 MI355X 可加入 vLLM CI,導致使用者體驗不佳。
Frame: NARY_RELATION · Mode: ASSERTED · Mapping: COMPLETE
開啟逐條審核

Atomic Claim 284/574 · 2026-02-16_inferencex-v2-nvidia-blackwell-vs::IX2-0299

Claim: 目前 vLLM 上針對 Mi355X 的 tests 數量為零。
Frame: ATTRIBUTE · Mode: ASSERTED · Mapping: COMPLETE
開啟逐條審核

Atomic Claim 285/574 · 2026-02-16_inferencex-v2-nvidia-blackwell-vs::IX2-0300

Claim: NVIDIA B200vLLM 上則已有大量 tests。
Frame: ATTRIBUTE · Mode: ASSERTED · Mapping: PARTIAL
開啟逐條審核

Atomic Claim 286/574 · 2026-02-16_inferencex-v2-nvidia-blackwell-vs::IX2-0301

Claim: 同樣地,vLLM 上目前也還沒有足夠的 MI300X CI machines。
Frame: ATTRIBUTE · Mode: ASSERTED · Mapping: PARTIAL
開啟逐條審核

Atomic Claim 287/574 · 2026-02-16_inferencex-v2-nvidia-blackwell-vs::IX2-0302

Claim: Upstream vLLM 至少還需要 20 台 MI300、20 台 MI325 與 20 台 MI355X machines,才能達到接近 CUDA 的可用性水準。
Frame: ATTRIBUTE · Mode: ASSERTED · Mapping: COMPLETE
開啟逐條審核

We at SemiAnalysis have been trying to get AMD to contribute more compute to vLLM and have had some success on that within the couple weeks. vLLM will start to get a couple of MI355X machines such that they can bring their CI test parity from 0% to non-0%. We will talk more about AMD’s previous lackluster contribution towards vLLM, SGLang, PyTorch CI machine situation & how Anush started to fix it in our upcoming State of AMD article. At SemiAnalysis, we will have internal dashboard to track the # of tests & quality of tests that AMD & NVIDIA runs on vLLM, SGLang, PyTorch, & JAX.

SemiAnalysis 過去一直在推 AMD 提供更多 compute 給 vLLM,最近幾週終於有一些成果。vLLM 會開始拿到幾台 MI355X,讓 CI test parity 從 0% 至少變成非零。之後《State of AMD》會更完整談 AMD 過去對 vLLM、SGLang、PyTorch CI machine 的貢獻為何偏弱,以及 Anush 開始如何改善。我們也會在 SemiAnalysis 建 internal dashboard,追蹤 AMD、NVIDIA 在 vLLM、SGLang、PyTorch、JAX 上的 test 數量與 test quality。

Atomic Claim 288/574 · 2026-02-16_inferencex-v2-nvidia-blackwell-vs::IX2-0303

Claim: SemiAnalysis 一直推動 AMDvLLM 提供更多算力,最近幾週已取得一些進展。
Frame: RELATION · Mode: ASSERTED · Mapping: COMPLETE
開啟逐條審核

Atomic Claim 289/574 · 2026-02-16_inferencex-v2-nvidia-blackwell-vs::IX2-0304

Claim: vLLM 將開始取得幾台 MI355X machines,使其 CI test parity 從 0% 提升到至少非零。
Frame: ATTRIBUTE · Mode: ASSERTED · Mapping: PARTIAL
開啟逐條審核

Atomic Claim 290/574 · 2026-02-16_inferencex-v2-nvidia-blackwell-vs::IX2-0305

Claim: SemiAnalysis 將在後續「State of AMD」文章中,更深入討論 AMD 過去對 vLLMSGLangPyTorch CI machines 貢獻不足,以及 Anush 如何開始改善這個問題。
Frame: ATTRIBUTE · Mode: ASSERTED · Mapping: COMPLETE
開啟逐條審核

Atomic Claim 291/574 · 2026-02-16_inferencex-v2-nvidia-blackwell-vs::IX2-0306

Claim: SemiAnalysis 將建立內部 dashboard,追蹤 AMDNVIDIAvLLMSGLangPyTorchJAX 上執行的 tests 數量與品質。
Frame: ATTRIBUTE · Mode: ASSERTED · Mapping: COMPLETE
開啟逐條審核

Moreover, the vLLM maintainers say that they cannot support day 0 vLLM support for ROCm due to this issue of lack of machine resources. This huge disparity in time to market continues to lead to ROCm lagging behind and leaving a huge opening for Nvidia to continue to charge an insane 75% gross margin (4x markup on cost of goods).

vLLM maintainer 也表示,正因為缺乏 machine resource,他們沒辦法替 ROCm 提供 Day 0 vLLM support。這種巨大的 time-to-market 差距會讓 ROCm 持續落後,也替 Nvidia 留下空間,繼續收取驚人的 75% gross margin——大約是 COGS 的 4x markup。

Atomic Claim 292/574 · 2026-02-16_inferencex-v2-nvidia-blackwell-vs::IX2-0307

Claim: 此外,vLLM maintainers 表示,因為缺乏 machine resources,目前無法為 ROCm 提供 day-0 vLLM support。
Frame: RELATION · Mode: ASSERTED · Mapping: COMPLETE
開啟逐條審核

Atomic Claim 293/574 · 2026-02-16_inferencex-v2-nvidia-blackwell-vs::IX2-0308

Claim: 這種 time-to-market 的巨大差距持續讓 ROCm 落後,也讓 Nvidia 有空間維持約 75% 的高毛利率,相當於對 COGS 約 4 倍加價。
Frame: COMPARISON · Mode: ASSERTED · Mapping: COMPLETE
開啟逐條審核

image

Source: Github

Lastly, AMD has not had enough committers “who demonstrated sustained upstream engagement through feature shepherding and code ownership” and has a lack of reviewers that can review their own code. This is why the pace of development on ROCm vLLM has been much slower than for CUDA vLLM.

最後,AMD 也缺少足夠的 committer——也就是那些能『透過 feature shepherding 與 code ownership 持續參與 upstream』的人——同時也缺乏可以 review 自家 code 的 reviewer。這就是為什麼 ROCm vLLM 的 development pace 一直比 CUDA vLLM 慢很多。

Atomic Claim 294/574 · 2026-02-16_inferencex-v2-nvidia-blackwell-vs::IX2-0309

Claim: 最後,AMD 也缺乏足夠多能持續參與 upstream、負責 feature shepherding 與 code ownership 的 committers。
Frame: ATTRIBUTE · Mode: ASSERTED · Mapping: COMPLETE
開啟逐條審核

Atomic Claim 295/574 · 2026-02-16_inferencex-v2-nvidia-blackwell-vs::IX2-0310

Claim: AMD 缺少足以 review 自家 code 的 reviewers。
Frame: ATTRIBUTE · Mode: ASSERTED · Mapping: PARTIAL
開啟逐條審核

Atomic Claim 296/574 · 2026-02-16_inferencex-v2-nvidia-blackwell-vs::IX2-0311

Claim: 這也是 ROCm vLLM 開發速度明顯慢於 CUDA vLLM 的原因之一。
Frame: ATTRIBUTE · Mode: ASSERTED · Mapping: PARTIAL
開啟逐條審核

There are many talented 10x engineers at AMD that work on ATOM and we would encourage AMD management to think about re-deploying these 10x engineers towards working on libraries and frameworks that people actually use, such as vLLM and SGLang.

AMD 其實有很多非常強的 10x engineer 在做 ATOM;我們鼓勵 AMD management 思考,把這些人才重新部署到真正有人使用的 library、framework,例如 vLLM 與 SGLang。

Atomic Claim 297/574 · 2026-02-16_inferencex-v2-nvidia-blackwell-vs::IX2-0312

Claim: AMD 有許多優秀工程師投入 ATOM,SemiAnalysis 建議 AMD management 考慮把這些人才轉向 vLLMSGLang 等實際被廣泛使用的 libraries 與 frameworks。
Frame: NARY_RELATION · Mode: ASSERTED · Mapping: COMPLETE
開啟逐條審核

As we mentioned earlier, AMD also needs to prioritize addressing composability issues with FP4, wideEP and disaggregated serving as opposed to overly focusing on optimizing FP4 for a single node.

如前面所說,相較於過度專注 single-node FP4 optimization,AMD 更應優先處理 FP4、wideEP、disaggregated serving 組合在一起時的 composability 問題。

Atomic Claim 298/574 · 2026-02-16_inferencex-v2-nvidia-blackwell-vs::IX2-0313

Claim: 如前文所述,AMD 也應優先解決 FP4、wideEP 與 disaggregated serving 的 composability 問題,而不是過度聚焦 single-node FP4 最佳化。
Frame: ATTRIBUTE · Mode: ASSERTED · Mapping: COMPLETE
開啟逐條審核

image

Source: SemiAnalysis InferenceX

Multi Token Prediction (MTP)

Speculative decoding reduces the cost of autoregressive generation by using a small, inexpensive draft model to propose several tokens ahead. The large model then checks the proposed tokens in a single forward pass that resembles a prefill computation. For a given input sequence length, a single forward pass can take roughly the same time when the input has N more tokens. Speculative decoding uses this property to run inference on a smaller model to draft multiple tokens for the main model to verify with a single forward pass, producing at most N additional tokens in a similar time budget.

Speculative decoding 的做法,是用一個較小、便宜的 draft model 先預測接下來多個 token,再讓大 model 用單次 forward pass 一次驗證,這個 verification 很像 prefill computation。對固定 input sequence length 而言,即使 input 多 N 個 token,單次 forward pass 所需時間可能仍大致相同。Speculative decoding 就利用這個特性,先讓小 model draft 多個 token,再交給 main model 以一次 forward pass 驗證,因此在相近 time budget 下,最多可以多產生 N 個 token。

Atomic Claim 299/574 · 2026-02-16_inferencex-v2-nvidia-blackwell-vs::IX2-0314

Claim: Speculative decoding 透過較小、成本較低的 draft model 預先提出多個 token,降低 autoregressive generation 成本。
Frame: COMPARISON · Mode: ASSERTED · Mapping: PARTIAL
開啟逐條審核

Atomic Claim 300/574 · 2026-02-16_inferencex-v2-nvidia-blackwell-vs::IX2-0315

Claim: 大模型之後只需透過一次類似 prefill 的 forward pass,驗證這些候選 token。
Frame: ATTRIBUTE · Mode: ASSERTED · Mapping: PARTIAL
開啟逐條審核

Atomic Claim 301/574 · 2026-02-16_inferencex-v2-nvidia-blackwell-vs::IX2-0316

Claim: 對固定輸入 sequence length 而言,即使輸入多 N 個 token,一次 forward pass 所需時間可能仍大致相同。
Frame: COMPARISON · Mode: ASSERTED · Mapping: PARTIAL
開啟逐條審核

Atomic Claim 302/574 · 2026-02-16_inferencex-v2-nvidia-blackwell-vs::IX2-0317

Claim: Speculative decoding 利用此特性,先以較小模型一次 draft 多個 tokens,再讓主模型用單次 forward pass 驗證,在近似相同時間預算內最多產生 N 個額外 token。
Frame: COMPARISON · Mode: ASSERTED · Mapping: PARTIAL
開啟逐條審核

image

Source: Brendan Bycroft

This assumption regarding additional token production with the same time budget is strongest for dense models because batched verification can reuse the same weight stream across multiple positions. For Mixture-of-Experts models, different tokens may route to different experts, so verifying multiple draft tokens can activate more experts than single-token decoding and force additional expert weights to be fetched from memory. As shown in the Mixtral 8x7B Instruct model results in the EAGLE paper, this extra memory traffic erodes bandwidth savings and can make verification notably comparable to a standard decoding step.

『相同 time budget 可以多產 token』這個假設在 dense model 最成立,因為 batched verification 可以讓多個 position 共用同一串 weight stream。MoE model 則不同:不同 token 可能 route 到不同 expert,因此一次驗證多個 draft token,可能比 single-token decode 啟用更多 expert,迫使 system 從 memory 額外抓更多 expert weight。EAGLE paper 的 Mixtral 8x7B Instruct result 就顯示,這些額外 memory traffic 會侵蝕 bandwidth saving,讓 verification cost 變得和一般 decode step 更接近。

Atomic Claim 303/574 · 2026-02-16_inferencex-v2-nvidia-blackwell-vs::IX2-0318

Claim: 「在相同時間內產生更多 token」的假設對 dense models 最成立,因為 batched verification 可以在多個 positions 重複利用相同 weight stream。
Frame: COMPARISON · Mode: HYPOTHETICAL · Mapping: PARTIAL
開啟逐條審核

Atomic Claim 304/574 · 2026-02-16_inferencex-v2-nvidia-blackwell-vs::IX2-0319

Claim:Mixture-of-Experts models,不同 token 可能 route 到不同 experts,因此一次驗證多個 draft tokens 可能啟動更多 experts,迫使系統從記憶體載入額外 expert weights。
Frame: ATTRIBUTE · Mode: HYPOTHETICAL · Mapping: COMPLETE
開啟逐條審核

Atomic Claim 305/574 · 2026-02-16_inferencex-v2-nvidia-blackwell-vs::IX2-0320

Claim: EAGLE 論文中的 Mixtral 8x7B Instruct 結果顯示,額外 memory traffic 會侵蝕 bandwidth savings,使 verification 成本可能接近一般 decoding step。
Frame: ATTRIBUTE · Mode: ASSERTED · Mapping: COMPLETE
開啟逐條審核

Multi-token prediction pursues similar benefits without requiring a separate draft model. Auxiliary prediction heads are added to the model architecture, so a single model can propose several future tokens from the same underlying representation. This improves distribution alignment because the proposals come from the same model that ultimately scores them. Multi-token prediction also avoids the operational complexity of serving an additional model while still enabling multi-token generation strategies but requires the MTP heads to be pretrained alongside the main model.

Multi-token prediction(MTP)追求類似好處,但不需要另外跑一個 draft model。它在 model architecture 中加入 auxiliary prediction head,讓同一個 model 能從相同 underlying representation 一次提出多個未來 token。因為 proposal 和最後負責 scoring 的是同一個 model,distribution alignment 更好;同時也省掉 serving 額外 model 的 operational complexity,仍能做 multi-token generation。不過代價是 MTP head 必須在 main model pretraining 階段就一起訓練。

Atomic Claim 306/574 · 2026-02-16_inferencex-v2-nvidia-blackwell-vs::IX2-0321

Claim: Multi-token prediction 追求類似效益,但不需要額外的 draft model。
Frame: ATTRIBUTE · Mode: ASSERTED · Mapping: PARTIAL
開啟逐條審核

Atomic Claim 307/574 · 2026-02-16_inferencex-v2-nvidia-blackwell-vs::IX2-0324

Claim: Multi-token prediction 不需要額外 serving 一個模型,因此可避免操作複雜度,同時支援 multi-token generation;但 MTP heads 必須與主模型一起預訓練。
Frame: ATTRIBUTE · Mode: ASSERTED · Mapping: COMPLETE
開啟逐條審核

image

Source: SemiAnalysis InferenceX

Across all SKUs, enabling MTP results in performance gains. By making use of the typically unused logits to verify the extra tokens, minimal compute overhead is added, saving extra expensive weight loads during decode.

所有 SKU 開啟 MTP 都能看到 performance gain。它利用原本通常沒有被充分使用的 logits 來驗證額外 token,只增加很少 compute overhead,卻能減少 decode 時昂貴的額外 weight load。

Atomic Claim 308/574 · 2026-02-16_inferencex-v2-nvidia-blackwell-vs::IX2-0325

Claim: 在所有 SKUs 上,啟用 MTP 都能帶來效能提升。
Frame: ATTRIBUTE · Mode: ASSERTED · Mapping: PARTIAL
開啟逐條審核

Atomic Claim 309/574 · 2026-02-16_inferencex-v2-nvidia-blackwell-vs::IX2-0326

Claim: MTP 利用通常未被使用的 logits 驗證額外 tokens,只增加很少 compute overhead,並可減少 decode 中昂貴的額外 weight loads。
Frame: ATTRIBUTE · Mode: ASSERTED · Mapping: PARTIAL
開啟逐條審核

image

Source: SemiAnalysis InferenceX

At large batch sizes, the inference regime is less memory-bandwidth bound compared to for low batch sizes. Since speculative decoding (including MTP) works by trading excess compute for fewer memory-bound decoding steps, this extra verification work from speculative tokens may not fit cleanly into slack, resulting in smaller improvements at high batch sizes.

Large batch size 時,inference 比 low batch size 更不容易受到 memory bandwidth 限制。Speculative decoding(包含 MTP)的核心是用多餘 compute 換取更少 memory-bound decode step;但 high batch 下本來就沒有那麼多 compute slack,因此 speculative token 的額外 verification work 未必能免費塞進空檔,performance improvement 也會比較小。

Atomic Claim 310/574 · 2026-02-16_inferencex-v2-nvidia-blackwell-vs::IX2-0327

Claim: 在大 batch size 下,推論較不受記憶體頻寬限制。
Frame: ATTRIBUTE · Mode: ASSERTED · Mapping: COMPLETE
開啟逐條審核

Atomic Claim 311/574 · 2026-02-16_inferencex-v2-nvidia-blackwell-vs::IX2-0328

Claim: 由於 speculative decoding(包含 MTP)本質上是用額外 compute 換取更少 memory-bound decoding steps,因此在高 batch size 下,speculative tokens 帶來的額外 verification 工作未必能完全塞進閒置 compute,效益可能較小。
Frame: COMPARISON · Mode: HYPOTHETICAL · Mapping: COMPLETE
開啟逐條審核

In terms of cost, MTP can drive huge cost savings, in the below table, we see that DeepSeek-R1-0528 run on FP4 using Dynamo TRT costs 0.057 per million total tokens.

從成本看,MTP 可以帶來非常大的 saving。下表中 DeepSeek-R1-0528 用 FP4 + Dynamo TRT 跑,每百萬 total token cost 約 $0.251;一旦開啟 MTP,成本可以大幅降到每百萬 total token 只有約 $0.057。

Atomic Claim 312/574 · 2026-02-16_inferencex-v2-nvidia-blackwell-vs::IX2-0329

Claim: 就成本而言,MTP 可帶來巨大節省;表中 DeepSeek-R1-0528 在 FP4 下使用 Dynamo TRT,每百萬 total tokens 成本約 0.251 美元。
Frame: RELATION · Mode: ASSERTED · Mapping: PARTIAL
開啟逐條審核

Atomic Claim 313/574 · 2026-02-16_inferencex-v2-nvidia-blackwell-vs::IX2-0330

Claim: 啟用 MTP 後,每百萬 total tokens 成本可大幅降至約 0.057 美元。
Frame: ATTRIBUTE · Mode: ASSERTED · Mapping: COMPLETE
開啟逐條審核

image

Source: SemiAnalysis InferenceX

In all configs, when all else is held equal, using MTP with DeepSeek R1 increases interactivity with no significant impact on model accuracy. This is in line with the DeepSeek V3 tech report findings.

所有 configuration 在其他條件相同時,DeepSeek R1 使用 MTP 都能提高 interactivity,而且對 model accuracy 沒有顯著影響。這和 DeepSeek V3 tech report 的發現一致。

Atomic Claim 314/574 · 2026-02-16_inferencex-v2-nvidia-blackwell-vs::IX2-0331

Claim: 在所有 configurations 中,其他條件相同時,DeepSeek R1 啟用 MTP 都能提高 interactivity,且不會對模型準確率造成顯著影響。
Frame: COMPARISON · Mode: ASSERTED · Mapping: COMPLETE
開啟逐條審核

Atomic Claim 315/574 · 2026-02-16_inferencex-v2-nvidia-blackwell-vs::IX2-0332

Claim: 這與 DeepSeek V3 tech report 的結果一致。
Frame: ATTRIBUTE · Mode: ASSERTED · Mapping: COMPLETE
開啟逐條審核

image

Source: SemiAnalysis InferenceX

Regarding the validity of MTP performance numbers, one may argue that the distribution of a synthetic dataset may not resemble real data. However, comparing MTP acceptance behavior between MTBench and our 1k1k benchmark, we see a very similar distribution confirming that our InferenceX benchmark is a good proxy for real world production performance. That said, InferenceX is not perfect and we are always looking to improve. If you want to be part of the mission, apply to join our special projects team here .

有人可能質疑 MTP performance number 的有效性,因為 synthetic dataset distribution 不一定像真實資料。但比較 MTBench 與我們 1k1k benchmark 的 MTP acceptance behavior,distribution 非常接近,證明 InferenceX benchmark 是 real-world production performance 的良好 proxy。當然 InferenceX 並不完美,我們也一直在改善。若想參與這項 mission,可以申請加入 special projects team ↗。

Atomic Claim 316/574 · 2026-02-16_inferencex-v2-nvidia-blackwell-vs::IX2-0333

Claim:MTP 效能數字的一項質疑是:synthetic dataset 的分布可能不像真實資料。
Frame: ATTRIBUTE · Mode: HYPOTHETICAL · Mapping: PARTIAL
開啟逐條審核

Atomic Claim 317/574 · 2026-02-16_inferencex-v2-nvidia-blackwell-vs::IX2-0334

Claim: 但比較 MTBench 與 InferenceX 1k1k benchmark 的 MTP acceptance behavior,可以看到非常相似的分布,支持 InferenceX benchmark 可作為 real-world production performance 的良好 proxy。
Frame: COMPARISON · Mode: ASSERTED · Mapping: COMPLETE
開啟逐條審核

Atomic Claim 318/574 · 2026-02-16_inferencex-v2-nvidia-blackwell-vs::IX2-0335

Claim: 不過 InferenceX 並不完美,團隊仍持續尋找改善方式。
Frame: ATTRIBUTE · Mode: ASSERTED · Mapping: PARTIAL
開啟逐條審核

Atomic Claim 319/574 · 2026-02-16_inferencex-v2-nvidia-blackwell-vs::IX2-0336

Claim: 原文此處邀請有興趣者加入其 special projects team。
Frame: CLAIM_ONLY · Mode: ASSERTED · Mapping: CLAIM_ONLY
開啟逐條審核

image

Source: SemiAnalysis InferenceX

Accuracy Evaluations

Throughput optimizations can sometimes quietly trade off accuracy (e.g. via aggressively relaxed acceptance rates, decoding tweaks, numerically unstable kernels, or endpoint misconfiguration). Without evals, a misconfigured server (truncation, bad decoding, wrong endpoint params) can still produce great throughput numbers but deliver garbage answers. For example, this additional layer of checks has helped us discover issues with some DP attention implementation for GPT-OSS.

Throughput optimization 有時會悄悄犧牲 accuracy,例如 acceptance rate 放得太寬、decoding tweak、numerically unstable kernel,或 endpoint configuration 錯誤。若沒有 eval,一台設定錯誤的 server——像 truncation、bad decoding、endpoint param 錯誤——仍可能跑出漂亮 throughput number,卻產生垃圾答案。這一層額外 check 就曾幫助我們找出 GPT-OSS 某些 DP attention implementation 的問題。

Atomic Claim 320/574 · 2026-02-16_inferencex-v2-nvidia-blackwell-vs::IX2-0337

Claim: 關於 InferenceX:Throughput optimizations 有時可能在不明顯的情況下犧牲 accuracy,例如過度放寬 acceptance rates、修改 decoding、使用數值不穩定 kernels,或 endpoint misconfiguration。
Frame: ATTRIBUTE · Mode: ASSERTED · Mapping: COMPLETE
開啟逐條審核

Atomic Claim 321/574 · 2026-02-16_inferencex-v2-nvidia-blackwell-vs::IX2-0338

Claim: 關於 InferenceX:若沒有 evals,即使 server configuration 錯誤,例如 truncation、bad decoding 或 endpoint params 錯誤,也可能得到漂亮 throughput 數字,但輸出品質非常差。
Frame: ATTRIBUTE · Mode: HYPOTHETICAL · Mapping: COMPLETE
開啟逐條審核

Atomic Claim 322/574 · 2026-02-16_inferencex-v2-nvidia-blackwell-vs::IX2-0339

Claim: 例如加入這層 accuracy checks 後,InferenceX 曾發現 GPT-OSS 某些 DP attention implementation 的問題。
Frame: ATTRIBUTE · Mode: ASSERTED · Mapping: PARTIAL
開啟逐條審核

Each representative throughput config now has an associated numerical accuracy check. Currently we are only using GSM8k, but being a very easy benchmark, the evaluation scores may not change much from differences in numerical calculation, and a harder benchmark may have a larger delta with respect to numerical accuracy. Thus, we plan to expand towards harder ones in the future, such as GPQA, HLE, MATH-500, SWE-Bench verified.

現在每個代表性 throughput configuration 都有對應 numerical accuracy check。目前只使用 GSM8k,但這個 benchmark 太容易,數值計算差異未必會明顯反映在 score;難度更高的 benchmark 對 numerical accuracy 的差異可能更敏感。因此未來我們會擴充到 GPQA、HLE、MATH-500、SWE-Bench Verified 等更困難 eval。

Atomic Claim 323/574 · 2026-02-16_inferencex-v2-nvidia-blackwell-vs::IX2-0340

Claim: 關於 InferenceX:現在每個代表性的 throughput configuration 都會搭配一項 numerical accuracy check。
Frame: ATTRIBUTE · Mode: ASSERTED · Mapping: COMPLETE
開啟逐條審核

Atomic Claim 324/574 · 2026-02-16_inferencex-v2-nvidia-blackwell-vs::IX2-0341

Claim: 關於 InferenceX:目前 accuracy check 只使用 GSM8k。
Frame: ATTRIBUTE · Mode: ASSERTED · Mapping: PARTIAL
開啟逐條審核

Atomic Claim 325/574 · 2026-02-16_inferencex-v2-nvidia-blackwell-vs::IX2-0342

Claim: 關於 InferenceX:由於 GSM8k 很容易,數值計算上的差異可能不會明顯反映在 evaluation score 上。
Frame: ATTRIBUTE · Mode: HYPOTHETICAL · Mapping: PARTIAL
開啟逐條審核

Atomic Claim 326/574 · 2026-02-16_inferencex-v2-nvidia-blackwell-vs::IX2-0343

Claim: 關於 InferenceX:更困難的 benchmark 在 numerical accuracy 上可能呈現更大的 delta。
Frame: ATTRIBUTE · Mode: HYPOTHETICAL · Mapping: PARTIAL
開啟逐條審核

Atomic Claim 327/574 · 2026-02-16_inferencex-v2-nvidia-blackwell-vs::IX2-0344

Claim: 因此未來計畫加入更難的 benchmarks,例如 GPQA、HLE、MATH-500 與 SWE-Bench verified
Frame: ATTRIBUTE · Mode: ASSERTED · Mapping: PARTIAL
開啟逐條審核

Another form of performance-accuracy tradeoff is quantization. Serving models at lower precision may result in worse model outputs. For DeepSeek R1, FP8 runs have very slightly higher evaluation scores than FP4. Note that GSM8k evals are saturated and often during QAT/PAT it is calibrated to common popular GSM8k, MATH-500, etc, leading to sometimes evals showing great results while real world end user evaluation being subpar. If we want to be part of the team to figure out how to properly evaluate inference engine accuracy, apply to join the mission here .

另一種 performance-accuracy trade-off 是 quantization。用較低 precision serving,可能讓 model output 變差。DeepSeek R1 的 FP8 run,evaluation score 就比 FP4 略高。要注意 GSM8k eval 已接近 saturation,而且 QAT/PAT 常會特別對 GSM8k、MATH-500 等熱門 benchmark calibration,結果可能是 benchmark score 很漂亮,但 real-world end-user evaluation 反而偏弱。若你也想一起研究 inference engine accuracy 該如何正確評估,可以申請加入這項 mission ↗。

Atomic Claim 328/574 · 2026-02-16_inferencex-v2-nvidia-blackwell-vs::IX2-0345

Claim: 另一種 performance-accuracy trade-off 是 quantization
Frame: ATTRIBUTE · Mode: ASSERTED · Mapping: PARTIAL
開啟逐條審核

Atomic Claim 329/574 · 2026-02-16_inferencex-v2-nvidia-blackwell-vs::IX2-0346

Claim: 關於 InferenceX:以較低 precision serving models 可能導致模型輸出品質變差。
Frame: COMPARISON · Mode: HYPOTHETICAL · Mapping: PARTIAL
開啟逐條審核

Atomic Claim 330/574 · 2026-02-16_inferencex-v2-nvidia-blackwell-vs::IX2-0347

Claim:DeepSeek R1 而言,FP8 runs 的 evaluation scores 略高於 FP4
Frame: COMPARISON · Mode: ASSERTED · Mapping: PARTIAL
開啟逐條審核

Atomic Claim 331/574 · 2026-02-16_inferencex-v2-nvidia-blackwell-vs::IX2-0348

Claim: 關於 InferenceX:需要注意,GSM8k eval 已高度飽和,而且 QAT/PAT 常針對 GSM8k、MATH-500 等熱門 benchmark 校準,因此 eval 看起來可能很好,但 real-world end-user evaluation 仍可能較差。
Frame: ATTRIBUTE · Mode: ASSERTED · Mapping: PARTIAL
開啟逐條審核

Atomic Claim 332/574 · 2026-02-16_inferencex-v2-nvidia-blackwell-vs::IX2-0349

Claim: 原文此處邀請有興趣者加入團隊,共同研究如何正確評估 inference engine accuracy。
Frame: ATTRIBUTE · Mode: ASSERTED · Mapping: PARTIAL
開啟逐條審核

image

Source: SemiAnalysis InferenceX

Anthropic Fast Mode Inferencing Explained

Anthropic recently released “fast mode ” alongside Opus 4.6. The value proposition: the same model quality at roughly 2.5× the speed, for around 6–12× the price. Both figures might seem surprising, and some users have speculated that this must require new hardware . It doesn’t. In fact, this is just the fundamental tradeoff at play. Any model can be served at a wide range of interactivity levels (tokens/sec per user), and the cost per million tokens (CPMT) shifts accordingly. Mercedes makes metro busses as well as race cars, to follow long with our analogy.

Anthropic 最近隨 Opus 4.6 推出『fast mode ↗』。Value proposition 很直接:model quality 相同,速度約 2.5x,但價格約 6–12x。這兩個數字看起來都很驚人,因此有人猜是不是用了新 hardware ↗;其實沒有。這只是 inference 最基本 trade-off 的結果。任何 model 都能在非常寬的 interactivity(tokens/sec per user)範圍 serving,而 cost per million tokens(CPMT)也會跟著移動。沿用前面的比喻,Mercedes 既可以造市區公車,也可以造賽車。

Atomic Claim 333/574 · 2026-02-16_inferencex-v2-nvidia-blackwell-vs::IX2-0350

Claim: Anthropic 最近隨 Opus 4.6 推出「fast mode」。
Frame: ATTRIBUTE · Mode: ASSERTED · Mapping: PARTIAL
開啟逐條審核

Atomic Claim 334/574 · 2026-02-16_inferencex-v2-nvidia-blackwell-vs::IX2-0351

Claim: 關於 Anthropic:其 value proposition 是:模型品質相同、速度約快 2.5 倍,但價格約高 6~12 倍。
Frame: COMPARISON · Mode: ASSERTED · Mapping: PARTIAL
開啟逐條審核

Atomic Claim 335/574 · 2026-02-16_inferencex-v2-nvidia-blackwell-vs::IX2-0352

Claim: 關於 Anthropic:這兩個數字看起來都可能令人意外。
Frame: ATTRIBUTE · Mode: HYPOTHETICAL · Mapping: COMPLETE
開啟逐條審核

Atomic Claim 336/574 · 2026-02-16_inferencex-v2-nvidia-blackwell-vs::IX2-0353

Claim: 關於 Anthropic:部分使用者猜測這一定需要新的硬體。
Frame: ATTRIBUTE · Mode: ASSERTED · Mapping: PARTIAL
開啟逐條審核

Atomic Claim 337/574 · 2026-02-16_inferencex-v2-nvidia-blackwell-vs::IX2-0354

Claim: 關於 Anthropic:實際上,這只是推論基本 trade-off 的自然結果。
Frame: ATTRIBUTE · Mode: ASSERTED · Mapping: PARTIAL
開啟逐條審核

Atomic Claim 338/574 · 2026-02-16_inferencex-v2-nvidia-blackwell-vs::IX2-0355

Claim: 關於 Anthropic:任何模型都可以在很寬的 interactivity(tokens/sec/user)範圍內 serving。
Frame: ATTRIBUTE · Mode: ASSERTED · Mapping: COMPLETE
開啟逐條審核

Atomic Claim 339/574 · 2026-02-16_inferencex-v2-nvidia-blackwell-vs::IX2-0356

Claim: 關於 Anthropic:每百萬 token 成本(CPMT)會隨 interactivity 改變。
Frame: ATTRIBUTE · Mode: ASSERTED · Mapping: COMPLETE
開啟逐條審核

Atomic Claim 340/574 · 2026-02-16_inferencex-v2-nvidia-blackwell-vs::IX2-0357

Claim: 沿用前述比喻,Mercedes 同時生產大眾巴士與賽車。
Frame: CLAIM_ONLY · Mode: ASSERTED · Mapping: CLAIM_ONLY
開啟逐條審核

Bean counters may think that fast mode is more expensive, but when looking at it through a total cost of ownership lens, fast mode is actually way cheaper for some situations. For example, a GB200 NVL72 rack can cost 3.3 million dollars, and as such, if claude code agentic loops (which runs on Trainium in production) that tool use call NVL72 racks, and these racks run inference 2.5x slower, you would need 2.5x more racks to deliver inference, meaning that not enabling fast mode would cost close to 5 million dollars in extra spend.

只看帳面價格,bean counter 可能覺得 fast mode 比較貴;但從 total cost of ownership 角度看,某些情況 fast mode 反而便宜很多。例如一個 GB200 NVL72 rack 約要 $3.3 million。假設 Claude Code agentic loop——production 實際跑在 Trainium——的 tool-use call 使用 NVL72 rack,而這些 rack 若 inference 慢 2.5x,就必須部署 2.5x 更多 rack 才能交付同樣 inference capacity。也就是說,不啟用 fast mode 反而可能多花接近 $5 million。

Atomic Claim 341/574 · 2026-02-16_inferencex-v2-nvidia-blackwell-vs::IX2-0358

Claim: 關於 Anthropic:只從單價看,財務人員可能認為 fast mode 更貴。
Frame: ATTRIBUTE · Mode: HYPOTHETICAL · Mapping: PARTIAL
開啟逐條審核

Atomic Claim 342/574 · 2026-02-16_inferencex-v2-nvidia-blackwell-vs::IX2-0359

Claim: 關於 Anthropic:但若從 total cost of ownership 角度看,fast mode 在某些情境其實反而便宜很多。
Frame: ATTRIBUTE · Mode: ASSERTED · Mapping: COMPLETE
開啟逐條審核

Atomic Claim 343/574 · 2026-02-16_inferencex-v2-nvidia-blackwell-vs::IX2-0360

Claim: 例如一個 GB200 NVL72 rack 成本約 330 萬美元。
Frame: ATTRIBUTE · Mode: ASSERTED · Mapping: COMPLETE
開啟逐條審核

Atomic Claim 344/574 · 2026-02-16_inferencex-v2-nvidia-blackwell-vs::IX2-0361

Claim: 假設 claude code agentic loops(production 實際運行於 Trainium)需要呼叫 NVL72 racks 執行 tool use。
Frame: ATTRIBUTE · Mode: HYPOTHETICAL · Mapping: COMPLETE
開啟逐條審核

Atomic Claim 345/574 · 2026-02-16_inferencex-v2-nvidia-blackwell-vs::IX2-0362

Claim: 關於 Claude Code:若這些 racks 的 inference 速度慢 2.5 倍,就需要 2.5 倍機櫃才能提供同等 inference capacity,因此不啟用 fast mode 可能需要多支出接近 500 萬美元。
Frame: COMPARISON · Mode: HYPOTHETICAL · Mapping: PARTIAL
開啟逐條審核

image

Source: Anthropic

image

Source: Anthropic

Consider a DeepSeek R1 0528 FP4 coding workflow served on B200s with TRT-LLM. At an interactivity of 50 tok/sec/user, inference cost is approximately 4/M output tokens, a 2.5× speed increase for a ~7× price increase, closely mirroring what we see with Anthropic’s fast mode. Note that this assumes DeepSeek R1 is similar to Opus 4.6, which isn’t the case. Still, the general principle holds true.

以 B200 + TRT-LLM 跑 DeepSeek R1 0528 FP4 coding workflow 為例:50 tok/sec/user 時,inference cost 約每百萬 output token $0.56;提高到 125 tok/sec/user 後,成本上升到約 $4/M output token。也就是速度提高 2.5x,價格約提高 7x,和 Anthropic fast mode 非常接近。當然這是假設 DeepSeek R1 和 Opus 4.6 類似,實際並不是;但背後 general principle 仍成立。

Atomic Claim 346/574 · 2026-02-16_inferencex-v2-nvidia-blackwell-vs::IX2-0363

Claim:B200 搭配 TRT-LLM serving DeepSeek R1 0528 FP4 coding workflow 為例。
Frame: ATTRIBUTE · Mode: ASSERTED · Mapping: COMPLETE
開啟逐條審核

Atomic Claim 347/574 · 2026-02-16_inferencex-v2-nvidia-blackwell-vs::IX2-0364

Claim: 關於 Anthropic:在 50 tok/sec/user interactivity 下,inference cost 約為每百萬 output tokens 0.56 美元。
Frame: ATTRIBUTE · Mode: ASSERTED · Mapping: COMPLETE
開啟逐條審核

Atomic Claim 348/574 · 2026-02-16_inferencex-v2-nvidia-blackwell-vs::IX2-0365

Claim: 在 125 tok/sec/user 下,成本升至約每百萬 output tokens 4 美元,也就是速度提高 2.5 倍、價格提高約 7 倍,與 Anthropic fast mode 的觀察非常接近。
Frame: COMPARISON · Mode: ASSERTED · Mapping: PARTIAL
開啟逐條審核

Atomic Claim 349/574 · 2026-02-16_inferencex-v2-nvidia-blackwell-vs::IX2-0366

Claim: 需注意,這裡假設 DeepSeek R1 與 Opus 4.6 類似,但實際並非如此。
Frame: ATTRIBUTE · Mode: HYPOTHETICAL · Mapping: PARTIAL
開啟逐條審核

Atomic Claim 350/574 · 2026-02-16_inferencex-v2-nvidia-blackwell-vs::IX2-0367

Claim: 關於 Anthropic:即使如此,背後的一般原理仍成立。
Frame: ATTRIBUTE · Mode: ASSERTED · Mapping: PARTIAL
開啟逐條審核

image

Source: SemiAnalysis InferenceX

image

Source: SemiAnalysis InferenceX

This follows directly from the fundamental latency-throughput tradeoff in LLM inference. At high batch sizes, GPUs achieve better utilization and greater total token throughput, meaning more users served concurrently and lower cost per token. At low batch sizes with greater parallelism per request, each user gets faster responses, but total token throughput drops. Since the hourly cost of the accelerators is fixed regardless of how they’re used, lower throughput means fewer tokens over which to amortize that cost, and thus a higher price per token.

這直接來自 LLM inference 最根本的 latency-throughput trade-off。High batch size 時 GPU utilization 更好,total token throughput 更高,可以同時服務更多 user,cost per token 也更低;low batch size、單一 request 分到更多 parallelism 時,每個 user response 更快,但 total token throughput 下降。Accelerator 每小時成本 ↗ 不會因使用方式不同而改變,因此 throughput 越低,可分攤 fixed cost 的 token 越少,price per token 自然越高。

Atomic Claim 351/574 · 2026-02-16_inferencex-v2-nvidia-blackwell-vs::IX2-0371

Claim: 關於 Anthropic:由於 accelerators 每小時成本不會因使用方式改變,throughput 越低,可分攤固定成本的 tokens 越少,因此每 token 價格越高。
Frame: ATTRIBUTE · Mode: ASSERTED · Mapping: PARTIAL
開啟逐條審核

Atomic Claim 352/574 · 2026-02-16_inferencex-v2-nvidia-blackwell-vs::IX2-0368

Claim: 這直接來自 LLM inference 的基本 latency-throughput trade-off。
Frame: ATTRIBUTE · Mode: ASSERTED · Mapping: COMPLETE
開啟逐條審核

Atomic Claim 353/574 · 2026-02-16_inferencex-v2-nvidia-blackwell-vs::IX2-0369

Claim: 在高 batch size 下,GPUs 利用率與總 token throughput 更高,因此可同時服務更多使用者,降低每 token 成本。
Frame: ATTRIBUTE · Mode: ASSERTED · Mapping: COMPLETE
開啟逐條審核

Atomic Claim 354/574 · 2026-02-16_inferencex-v2-nvidia-blackwell-vs::IX2-0370

Claim: 在低 batch size、每個 request 使用更多 parallelism 時,每位使用者回應會更快,但總 token throughput 下降。
Frame: ATTRIBUTE · Mode: EXPECTED · Mapping: COMPLETE
開啟逐條審核

In short, fast mode isn’t necessarily a hardware story, but merely the natural consequence of trading throughput for latency on the same GPUs.

簡單說,fast mode 不一定是 hardware story;它只是同一批 GPU 上,用 throughput 換 latency 的自然結果。

Atomic Claim 355/574 · 2026-02-16_inferencex-v2-nvidia-blackwell-vs::IX2-0372

Claim: 因此 fast mode 不一定是硬體升級,而可能只是同一批 GPUs 上用 throughput 換 latency 的自然結果。
Frame: ATTRIBUTE · Mode: ASSERTED · Mapping: COMPLETE
開啟逐條審核

image

Source: SemiAnalysis InferenceX

Furthermore, we observe that inference optimization techniques such as speculative decoding, as explained earlier, can directly lead to cheaper inference; no new chips are required.

此外,如前面解釋的 speculative decoding 等 inference optimization,本身就能直接降低 inference cost,完全不需要換新 chip。

Atomic Claim 356/574 · 2026-02-16_inferencex-v2-nvidia-blackwell-vs::IX2-0373

Claim: 此外,前文所述的 speculative decoding 等 inference optimization techniques,可直接降低 inference 成本。
Frame: ATTRIBUTE · Mode: ASSERTED · Mapping: COMPLETE
開啟逐條審核

Atomic Claim 357/574 · 2026-02-16_inferencex-v2-nvidia-blackwell-vs::IX2-0374

Claim: 關於 speculative decoding:這些改善不需要新的晶片。
Frame: ATTRIBUTE · Mode: ASSERTED · Mapping: PARTIAL
開啟逐條審核

Take the following example, DeepSeek R1 FP4 on an 8k/1k workload. At an interactivity level of 150 tok/sec/user, the baseline GB300 Dynamo TRT cost per million tokens is approximately 0.11. This is a ~21x price decrease at this interactivity level simply by employing an inference optimization technique.

以 DeepSeek R1 FP4、8k/1k workload 為例。在 150 tok/sec/user interactivity 下,baseline GB300 Dynamo TRT 每百萬 token cost 約 $2.35;一旦開啟 MTP,價格可降到約 $0.11。也就是只靠 inference optimization,同一 interactivity 下成本直接下降約 21x。

Atomic Claim 358/574 · 2026-02-16_inferencex-v2-nvidia-blackwell-vs::IX2-0375

Claim: 以下以 8k/1k workload 上的 DeepSeek R1 FP4 為例。
Frame: ATTRIBUTE · Mode: ASSERTED · Mapping: PARTIAL
開啟逐條審核

Atomic Claim 359/574 · 2026-02-16_inferencex-v2-nvidia-blackwell-vs::IX2-0376

Claim: 在 150 tok/sec/user interactivity 下,baseline GB300 Dynamo TRT 每百萬 token 成本約 2.35 美元。
Frame: ATTRIBUTE · Mode: ASSERTED · Mapping: COMPLETE
開啟逐條審核

Atomic Claim 360/574 · 2026-02-16_inferencex-v2-nvidia-blackwell-vs::IX2-0377

Claim: 啟用 MTP 後,價格可降至約 0.11 美元。
Frame: ATTRIBUTE · Mode: ASSERTED · Mapping: COMPLETE
開啟逐條審核

Atomic Claim 361/574 · 2026-02-16_inferencex-v2-nvidia-blackwell-vs::IX2-0378

Claim: 關於 Anthropic:在相同 interactivity 下,只靠 inference optimization 就可讓價格下降約 21 倍。
Frame: COMPARISON · Mode: ASSERTED · Mapping: PARTIAL
開啟逐條審核

image

Source: SemiAnalysis InferenceX

image

Source: SemiAnalysis InferenceX

image

Source: SemiAnalysis InferenceX

Fixing an interactivity level of 50 tok/sec/user, we further see how much MTP can effectively decrease CPMT across a variety of chips.

若固定 interactivity = 50 tok/sec/user,也能看到 MTP 在不同 chip 上可以把 CPMT 壓低多少。

Atomic Claim 362/574 · 2026-02-16_inferencex-v2-nvidia-blackwell-vs::IX2-0379

Claim: 固定 interactivity 50 tok/sec/user 後,也能看到 MTP 在各種晶片上都可有效降低 CPMT。
Frame: ATTRIBUTE · Mode: ASSERTED · Mapping: COMPLETE
開啟逐條審核

image

Source: SemiAnalysis InferenceX

Wide Expert Parallelism (WideEP) and Disaggregated Prefill

In this section, we will go deeper on expert parallelism and go on to explain what _wide _expert parallelism is. We will then explain the idea of Disaggregated Prefill, how it is different from WideEP, and how WideEP and Disaggregated Prefill are used in unison to achieve SOTA performance.

這一節會更深入談 expert parallelism,接著說明什麼是 wide expert parallelism。我們也會解釋 Disaggregated Prefill 的概念、它和 WideEP 有什麼不同,以及 WideEP + Disaggregated Prefill 如何搭配使用,做到 SOTA performance。

Atomic Claim 363/574 · 2026-02-16_inferencex-v2-nvidia-blackwell-vs::IX2-0380

Claim: 本節會更深入討論 expert parallelism,並說明 wide expert parallelism 的概念。
Frame: ATTRIBUTE · Mode: ASSERTED · Mapping: PARTIAL
開啟逐條審核

Atomic Claim 364/574 · 2026-02-16_inferencex-v2-nvidia-blackwell-vs::IX2-0381

Claim: 接著會介紹 Disaggregated Prefill,以及它與 WideEP 的差異。
Frame: ATTRIBUTE · Mode: ASSERTED · Mapping: PARTIAL
開啟逐條審核

Atomic Claim 365/574 · 2026-02-16_inferencex-v2-nvidia-blackwell-vs::IX2-0382

Claim: 並說明 WideEP 與 Disaggregated Prefill 如何搭配使用以達成 SOTA performance。
Frame: ATTRIBUTE · Mode: ASSERTED · Mapping: PARTIAL
開啟逐條審核

WideEP

By now, most frontier AI labs employ Mixture of Experts (MoE) model architectures as opposed to dense. In MoE architectures, only a subset of “experts” are activated for each token. For instance, DeepSeek R1 has 671B total parameters, but only 37B active parameters. Specifically, DeepSeek R1 has 256 routed experts (and 1 shared expert) with each token being routed to 8 distinct experts. This architecture lends itself naturally to expert parallelism (EP), which evenly distributes expert weights across some number of GPUs.

現在多數 frontier AI lab 已經採 Mixture of Experts(MoE)architecture,而不是 dense model。MoE 每個 token 只啟用一部分 expert。例如 DeepSeek R1 總參數 671B,但 active parameter 只有 37B;具體來說,它有 256 個 routed expert(再加 1 個 shared expert),每個 token 會 route 到 8 個不同 expert。這種 architecture 天然適合 expert parallelism(EP),把 expert weight 平均分散到多顆 GPU。

Atomic Claim 366/574 · 2026-02-16_inferencex-v2-nvidia-blackwell-vs::IX2-0383

Claim: 目前大多數 frontier AI labs 都採用 Mixture of Experts (MoE) architecture,而不是 dense model。
Frame: ATTRIBUTE · Mode: ASSERTED · Mapping: COMPLETE
開啟逐條審核

Atomic Claim 367/574 · 2026-02-16_inferencex-v2-nvidia-blackwell-vs::IX2-0384

Claim:MoE architecture 中,每個 token 只會啟動部分「experts」。
Frame: ATTRIBUTE · Mode: ASSERTED · Mapping: COMPLETE
開啟逐條審核

Atomic Claim 368/574 · 2026-02-16_inferencex-v2-nvidia-blackwell-vs::IX2-0385

Claim: 例如 DeepSeek R1 有 671B total parameters,但 active parameters 只有 37B。
Frame: ATTRIBUTE · Mode: ASSERTED · Mapping: PARTIAL
開啟逐條審核

Atomic Claim 369/574 · 2026-02-16_inferencex-v2-nvidia-blackwell-vs::IX2-0386

Claim: 更具體而言,DeepSeek R1 有 256 個 routed experts(另有 1 個 shared expert),每個 token 會被 route 到 8 個不同 experts。
Frame: ATTRIBUTE · Mode: ASSERTED · Mapping: COMPLETE
開啟逐條審核

Atomic Claim 370/574 · 2026-02-16_inferencex-v2-nvidia-blackwell-vs::IX2-0387

Claim: 這種架構天然適合 expert parallelism (EP),也就是把 expert weights 平均分配到多顆 GPUs
Frame: ATTRIBUTE · Mode: ASSERTED · Mapping: COMPLETE
開啟逐條審核

Consider serving DeepSeek R1 on a single 8-GPU server. At 671B parameters, some form of parallelism is required to fit the model across available HBM. The naive approach is tensor parallelism (TP), which shards every weight matrix across all GPUs. This works well for dense models but ignores the sparse activation pattern of MoE. With TP=8, each expert’s weights are sharded across all 8 GPUs, meaning every expert activation requires an all-reduce across all GPUs & the reduction dims of the GEMM is smaller leading to lower arithmetic intensity, even though only 8 of 256 experts activate per token. TP treats each expert like a dense layer, paying full cross-GPU communication cost while the model’s sparsity goes unexploited.

假設在單一 8-GPU server 上 serving DeepSeek R1。671B parameter 不可能塞進單顆 GPU,因此一定要做某種 parallelism。最直覺是 tensor parallelism(TP),把每個 weight matrix 都 shard 到 8 顆 GPU。這對 dense model 很合理,但完全忽略 MoE sparse activation。TP=8 時,每個 expert weight 也被切到全部 8 顆 GPU,所以每次 expert activation 都得跨 8 GPU 做 all-reduce;同時 GEMM reduction dimension 變小、arithmetic intensity 降低。即使每個 token 實際只啟用 256 個 expert 中的 8 個,TP 還是把每個 expert 當 dense layer 處理,支付完整 cross-GPU communication cost,沒有吃到 sparsity 優勢。

Atomic Claim 371/574 · 2026-02-16_inferencex-v2-nvidia-blackwell-vs::IX2-0388

Claim: 可以考慮在單一 8-GPU server 上 serving DeepSeek R1。
Frame: ATTRIBUTE · Mode: ASSERTED · Mapping: PARTIAL
開啟逐條審核

Atomic Claim 372/574 · 2026-02-16_inferencex-v2-nvidia-blackwell-vs::IX2-0389

Claim: 模型有 671B parameters,因此必須使用某種 parallelism,才能把模型放進可用 HBM
Frame: ATTRIBUTE · Mode: ASSERTED · Mapping: COMPLETE
開啟逐條審核

Atomic Claim 373/574 · 2026-02-16_inferencex-v2-nvidia-blackwell-vs::IX2-0390

Claim: 最直接的方法是 tensor parallelism (TP),把每個 weight matrix 分片到所有 GPUs
Frame: RELATION · Mode: ASSERTED · Mapping: COMPLETE
開啟逐條審核

Atomic Claim 374/574 · 2026-02-16_inferencex-v2-nvidia-blackwell-vs::IX2-0391

Claim: 這對 dense models 很有效,但忽略 MoE 的 sparse activation pattern。
Frame: ATTRIBUTE · Mode: ASSERTED · Mapping: PARTIAL
開啟逐條審核

Atomic Claim 375/574 · 2026-02-16_inferencex-v2-nvidia-blackwell-vs::IX2-0392

Claim: 在 TP=8 下,每個 expert weights 都分散在 8 顆 GPUs,因此每次 expert activation 都需要跨所有 GPUs 執行 all-reduce;同時 GEMM reduction dimension 變小,降低 arithmetic intensity,即使每 token 實際只啟動 256 個 experts 中的 8 個。
Frame: ATTRIBUTE · Mode: ASSERTED · Mapping: COMPLETE
開啟逐條審核

Atomic Claim 376/574 · 2026-02-16_inferencex-v2-nvidia-blackwell-vs::IX2-0393

Claim: TP 把每個 expert 當成 dense layer,支付完整 cross-GPU communication 成本,而沒有利用模型的 sparsity
Frame: ATTRIBUTE · Mode: ASSERTED · Mapping: COMPLETE
開啟逐條審核

Expert parallelism takes a more well-suited approach, assigning whole experts to individual GPUs. With EP=8, we divide the 256 experts per layer across 8 GPUs for a total of 32 experts/layer/GPU. Each GPU holds approximately 1/8th of the expert weights plus a full replica of the non-expert weights (attention projections, embeddings, normalization, and the shared expert). Since roughly 90%+ of DeepSeek R1’s parameters are routed expert weights, EP captures most of the memory savings, and replicating the remaining less than 30B non-expert parameters across all 8 GPUs is affordable.

Expert parallelism 更貼合 MoE 特性:直接把完整 expert 分配給個別 GPU。EP=8 時,每層 256 個 expert 平均分到 8 顆 GPU,也就是每 GPU 每層 32 個 expert。每顆 GPU 持有大約 1/8 expert weight,再加一份完整 non-expert weight replica,包括 attention projection、embedding、normalization、shared expert。DeepSeek R1 超過 90% parameter 都是 routed expert weight,因此 EP 已經抓到絕大部分 memory saving;剩下不到 30B non-expert parameter 在 8 顆 GPU 各複製一份,成本仍可接受。

Atomic Claim 377/574 · 2026-02-16_inferencex-v2-nvidia-blackwell-vs::IX2-0394

Claim: Expert parallelism 更適合這種架構,它直接把完整 expert 分配給個別 GPUs
Frame: ATTRIBUTE · Mode: ASSERTED · Mapping: COMPLETE
開啟逐條審核

Atomic Claim 378/574 · 2026-02-16_inferencex-v2-nvidia-blackwell-vs::IX2-0395

Claim: 在 EP=8 下,每層 256 個 experts 平均分給 8 顆 GPUs,也就是每顆 GPU 每層持有 32 個 experts。
Frame: ATTRIBUTE · Mode: ASSERTED · Mapping: COMPLETE
開啟逐條審核

Atomic Claim 379/574 · 2026-02-16_inferencex-v2-nvidia-blackwell-vs::IX2-0396

Claim: 每顆 GPU 約持有 1/8 的 expert weights,外加完整複製一份 non-expert weights,例如 attention projections、embeddings、normalization 與 shared expert。
Frame: ATTRIBUTE · Mode: ASSERTED · Mapping: COMPLETE
開啟逐條審核

Atomic Claim 380/574 · 2026-02-16_inferencex-v2-nvidia-blackwell-vs::IX2-0397

Claim: 由於 DeepSeek R1 超過約 90% parameters 都是 routed expert weights,EP 可取得大部分記憶體節省效果。
Frame: ATTRIBUTE · Mode: ASSERTED · Mapping: PARTIAL
開啟逐條審核

Atomic Claim 381/574 · 2026-02-16_inferencex-v2-nvidia-blackwell-vs::IX2-0398

Claim: 剩下不到 30B 的 non-expert parameters 即使在 8 顆 GPUs 全部複製,成本仍可接受。
Frame: ATTRIBUTE · Mode: ASSERTED · Mapping: COMPLETE
開啟逐條審核

The forward pass proceeds in two phases per layer. During attention, each GPU acts as an independent data-parallel rank, processing its own subset of requests using its replicated non-expert weights, no inter-GPU communication is needed. During the MoE phase, a lightweight router determines which experts each token requires, and tokens are dispatched to the appropriate GPUs via all-to-all communication. Each GPU executes its local experts on only the tokens routed to it, and results are returned via a second all-to-all.

每層 forward pass 分兩個 phase。Attention phase 時,每顆 GPU 都像獨立 data-parallel rank,用自己完整複製的 non-expert weight 處理自己的 request subset,不需要 inter-GPU communication。MoE phase 時,輕量 router 決定每個 token 需要哪些 expert,再透過 all-to-all 把 token dispatch 到持有對應 expert 的 GPU;每顆 GPU 只對被 route 過來的 token 執行 local expert,結果再透過第二次 all-to-all 傳回。

Atomic Claim 382/574 · 2026-02-16_inferencex-v2-nvidia-blackwell-vs::IX2-0399

Claim: 關於 EP:每一層的 forward pass 分成兩個階段。
Frame: ATTRIBUTE · Mode: ASSERTED · Mapping: COMPLETE
開啟逐條審核

Atomic Claim 383/574 · 2026-02-16_inferencex-v2-nvidia-blackwell-vs::IX2-0400

Claim: 在 attention 階段,每顆 GPU 都作為獨立 data-parallel rank,利用已複製的 non-expert weights 處理自己的 request subset,因此不需要 inter-GPU communication。
Frame: ATTRIBUTE · Mode: ASSERTED · Mapping: COMPLETE
開啟逐條審核

Atomic Claim 384/574 · 2026-02-16_inferencex-v2-nvidia-blackwell-vs::IX2-0401

Claim:MoE 階段,一個輕量 router 會判斷每個 token 需要哪些 experts。
Frame: ATTRIBUTE · Mode: ASSERTED · Mapping: COMPLETE
開啟逐條審核

Atomic Claim 385/574 · 2026-02-16_inferencex-v2-nvidia-blackwell-vs::IX2-0402

Claim: 接著 tokens 透過 all-to-all communication 被送往對應的 GPUs
Frame: ATTRIBUTE · Mode: ASSERTED · Mapping: PARTIAL
開啟逐條審核

Atomic Claim 386/574 · 2026-02-16_inferencex-v2-nvidia-blackwell-vs::IX2-0403

Claim: 每顆 GPU 只會對被 routing 到本機的 tokens 執行其 local experts。
Frame: ATTRIBUTE · Mode: ASSERTED · Mapping: COMPLETE
開啟逐條審核

Atomic Claim 387/574 · 2026-02-16_inferencex-v2-nvidia-blackwell-vs::IX2-0404

Claim: 關於 EP:運算結果再透過第二次 all-to-all 傳回。
Frame: ATTRIBUTE · Mode: ASSERTED · Mapping: PARTIAL
開啟逐條審核

image

_An EP8 DP8 deployment of DeepSeek R1. All 256 experts per layer are divided evenly among the 8 GPUs, whereas attention along with other non-expert weights (shared expert, gating network, RMSNorm, LM head, etc.) are replicated across all 8 DP ranks. _Source: SemiAnalysis

DeepSeek R1 的 EP8 DP8 deployment。每層 256 個 expert 平均分散到 8 顆 GPU;attention 與其他 non-expert weight(shared expert、gating network、RMSNorm、LM head 等)則在 8 個 DP rank 全部複製。Source: SemiAnalysis

Atomic Claim 388/574 · 2026-02-16_inferencex-v2-nvidia-blackwell-vs::IX2-0405

Claim: 這是一個 DeepSeek R1 的 EP8 DP8 deployment。
Frame: ATTRIBUTE · Mode: ASSERTED · Mapping: COMPLETE
開啟逐條審核

Atomic Claim 389/574 · 2026-02-16_inferencex-v2-nvidia-blackwell-vs::IX2-0406

Claim: 每層 256 個 experts 平均分配到 8 顆 GPUs;attention 與其他 non-expert weights(shared expert、gating network、RMSNorm、LM head 等)則採複製方式。
Frame: ATTRIBUTE · Mode: ASSERTED · Mapping: COMPLETE
開啟逐條審核

The obvious way to scale is replication: deploy N independent EP8 instances across N nodes. Each instance serves requests independently with no cross-node communication. This scales throughput linearly, but each GPU still holds 32 experts per layer, and each token activates at most 8 of those 32 local experts. 75% of expert weights sit cold in HBM.

最直覺的 scale 方法是 replication:在 N 個 node 部署 N 個獨立 EP8 instance,各自 serving request,不做 cross-node communication。Throughput 可以線性 scale,但每顆 GPU 仍持有每層 32 個 expert,而一個 token 最多只會啟用其中 8 個 local expert,因此 75% expert weight 都只是冷冷躺在 HBM 裡。

Atomic Claim 390/574 · 2026-02-16_inferencex-v2-nvidia-blackwell-vs::IX2-0408

Claim: 關於 EP:最直觀的 scale-out 方法是 replication:在 N 個 nodes 上部署 N 組彼此獨立的 EP8 instances。
Frame: ATTRIBUTE · Mode: ASSERTED · Mapping: COMPLETE
開啟逐條審核

Atomic Claim 391/574 · 2026-02-16_inferencex-v2-nvidia-blackwell-vs::IX2-0409

Claim: 關於 EP:每個 instance 各自獨立服務 requests,不需要 cross-node communication。
Frame: ATTRIBUTE · Mode: ASSERTED · Mapping: COMPLETE
開啟逐條審核

Atomic Claim 392/574 · 2026-02-16_inferencex-v2-nvidia-blackwell-vs::IX2-0410

Claim: 這可讓 throughput 線性擴展,但每顆 GPU 仍需持有每層 32 個 experts,而每個 token 最多只會啟動其中 8 個 local experts。
Frame: ATTRIBUTE · Mode: ASSERTED · Mapping: COMPLETE
開啟逐條審核

Atomic Claim 393/574 · 2026-02-16_inferencex-v2-nvidia-blackwell-vs::IX2-0411

Claim: 因此約 75% 的 expert weights 會留在 HBM 中而沒有被使用。
Frame: ATTRIBUTE · Mode: ASSERTED · Mapping: PARTIAL
開啟逐條審核

Wide expert parallelism (WideEP) takes a different approach by scaling EP _across _nodes rather than replicating independent instances. On a 64-GPU cluster (8 nodes), DP64/EP64 places only 256/64 = 4 experts per layer per GPU, each still holding a full replica of the non-expert weights. During the MoE phase, tokens from all 64 DP ranks are dispatched via all-to-all to the GPUs hosting their routed experts.

Wide expert parallelism(WideEP)走另一條路:不是複製獨立 instance,而是把 EP 跨 node 擴大。以 64-GPU cluster(8 nodes)為例,DP64/EP64 讓每層 256 個 expert 平均分到 64 顆 GPU,因此每顆只放 4 個 expert;non-expert weight 仍是每顆 GPU 一份完整 replica。MoE phase 時,64 個 DP rank 的 token 會透過 all-to-all,dispatch 到持有相應 routed expert 的 GPU。

Atomic Claim 394/574 · 2026-02-16_inferencex-v2-nvidia-blackwell-vs::IX2-0412

Claim: Wide expert parallelism(WideEP)採用不同方式:不是複製獨立 instances,而是把 EP 直接跨 nodes 擴展。
Frame: ATTRIBUTE · Mode: ASSERTED · Mapping: COMPLETE
開啟逐條審核

Atomic Claim 395/574 · 2026-02-16_inferencex-v2-nvidia-blackwell-vs::IX2-0413

Claim: 在 64-GPU cluster(8 nodes)上,DP64/EP64 會讓每顆 GPU 每層只持有 256/64=4 個 experts,但仍完整複製 non-expert weights。
Frame: ATTRIBUTE · Mode: ASSERTED · Mapping: COMPLETE
開啟逐條審核

Atomic Claim 396/574 · 2026-02-16_inferencex-v2-nvidia-blackwell-vs::IX2-0414

Claim:MoE 階段,來自 64 個 DP ranks 的 tokens 會透過 all-to-all 被送到持有其 routed experts 的 GPUs
Frame: ATTRIBUTE · Mode: ASSERTED · Mapping: COMPLETE
開啟逐條審核

This yields three compounding benefits over the single-node EP8 baseline. First, reducing expert footprint from 32 to 4 experts/GPU frees substantial HBM for KV cache, directly increasing per-GPU batch size capacity. Second, 64 DP ranks funneling tokens through fewer experts per GPU increases tokens-per-expert, raising arithmetic intensity (more FLOPs per byte of weights loaded) and improving compute utilization. The same expert weights service 8x more tokens per step. Third, aggregate HBM bandwidth scales linearly with GPU count; 64 GPUs loading expert weights simultaneously provide 8x the memory bandwidth of a single node, reducing memory bottleneck.

相較 single-node EP8 baseline,這會帶來三個疊加優勢。第一,expert footprint 從每 GPU 32 個降到 4 個,釋放大量 HBM 給 KV cache,直接提高 per-GPU batch-size capacity。第二,64 個 DP rank 的 token 匯聚到每 GPU 更少的 expert,使 tokens-per-expert 上升,arithmetic intensity 提高,也就是每載入一 byte weight 能做更多 FLOPs,compute utilization 更好;同一份 expert weight 每 step 可服務 8x 更多 token。第三,aggregate HBM bandwidth 隨 GPU 數線性增加;64 顆 GPU 同時載 expert weight,總 memory bandwidth 是單 node 的 8x,memory bottleneck 因而大幅降低。

Atomic Claim 397/574 · 2026-02-16_inferencex-v2-nvidia-blackwell-vs::IX2-0415

Claim: 關於 EP:相較 single-node EP8 baseline,這會產生三項彼此疊加的優勢。
Frame: COMPARISON · Mode: ASSERTED · Mapping: PARTIAL
開啟逐條審核

Atomic Claim 398/574 · 2026-02-16_inferencex-v2-nvidia-blackwell-vs::IX2-0416

Claim: 第一,把每顆 GPU 的 expert footprint 從 32 個降到 4 個,可釋放大量 HBMKV cache,直接提高每顆 GPU 可承載的 batch size
Frame: ATTRIBUTE · Mode: ASSERTED · Mapping: COMPLETE
開啟逐條審核

Atomic Claim 399/574 · 2026-02-16_inferencex-v2-nvidia-blackwell-vs::IX2-0417

Claim: 第二,64 個 DP ranks 的 tokens 被集中到每顆 GPU 較少的 experts 上,使每個 expert 接收到更多 tokens,提高 arithmetic intensity,也就是每載入一 byte weights 可執行更多 FLOPs,進而提升 compute utilization。
Frame: COMPARISON · Mode: ASSERTED · Mapping: COMPLETE
開啟逐條審核

Atomic Claim 400/574 · 2026-02-16_inferencex-v2-nvidia-blackwell-vs::IX2-0418

Claim: 關於 EP:同一組 expert weights 在每個 step 可服務約 8 倍 tokens。
Frame: COMPARISON · Mode: ASSERTED · Mapping: PARTIAL
開啟逐條審核

Atomic Claim 401/574 · 2026-02-16_inferencex-v2-nvidia-blackwell-vs::IX2-0419

Claim: 第三,aggregate HBM bandwidth 會隨 GPU 數量近似線性增加。
Frame: ATTRIBUTE · Mode: ASSERTED · Mapping: COMPLETE
開啟逐條審核

Atomic Claim 402/574 · 2026-02-16_inferencex-v2-nvidia-blackwell-vs::IX2-0420

Claim: 64 顆 GPUs 同時載入 expert weights,可提供單一 node 約 8 倍的 memory bandwidth,降低記憶體瓶頸。
Frame: COMPARISON · Mode: ASSERTED · Mapping: PARTIAL
開啟逐條審核

image

_A WideEP EP64 DP64 deployment of DeepSeek R1. All 256 experts per layer are divided evenly among the 64 GPUs (8 nodes), and attention and other non-expert weights (shared expert, gating network, RMSNorm, LM head, etc.) are replicated across all 64 DP ranks. _Source: SemiAnalysis

DeepSeek R1 的 WideEP EP64 DP64 deployment。每層 256 個 expert 平均分散到 64 顆 GPU(8 nodes);attention 與其他 non-expert weight(shared expert、gating network、RMSNorm、LM head 等)則在全部 64 個 DP rank 複製。Source: SemiAnalysis

Atomic Claim 403/574 · 2026-02-16_inferencex-v2-nvidia-blackwell-vs::IX2-0421

Claim: 這是一個 DeepSeek R1 的 WideEP EP64 DP64 deployment。
Frame: ATTRIBUTE · Mode: ASSERTED · Mapping: COMPLETE
開啟逐條審核

Atomic Claim 404/574 · 2026-02-16_inferencex-v2-nvidia-blackwell-vs::IX2-0422

Claim: 每層 256 個 experts 平均分到 64 顆 GPUs(8 nodes),attention 與其他 non-expert weights(shared expert、gating network、RMSNorm、LM head 等)則採複製方式。
Frame: ATTRIBUTE · Mode: ASSERTED · Mapping: COMPLETE
開啟逐條審核

The above configurations use only DP+EP (also known as DEP), where each GPU holds a full replica of all non-expert weights. As GPU count grows, this replication becomes increasingly wasteful. On a 64-GPU DP64/EP64 deployment, every GPU stores an identical copy of the ~40B non-expert parameters.

上面的 configuration 只用 DP+EP,也稱 DEP,因此每顆 GPU 都持有一份完整 non-expert weight。GPU 數越多,這種 replication 越浪費。64-GPU DP64/EP64 deployment 中,每一顆 GPU 都存著完全相同的一份約 40B non-expert parameter。

Atomic Claim 405/574 · 2026-02-16_inferencex-v2-nvidia-blackwell-vs::IX2-0424

Claim: 上述 configurations 只使用 DP+EP(亦稱 DEP),因此每顆 GPU 都完整複製所有 non-expert weights。
Frame: ATTRIBUTE · Mode: ASSERTED · Mapping: COMPLETE
開啟逐條審核

Atomic Claim 406/574 · 2026-02-16_inferencex-v2-nvidia-blackwell-vs::IX2-0425

Claim:GPU 數量增加,這種 replication 會越來越浪費記憶體。
Frame: ATTRIBUTE · Mode: ASSERTED · Mapping: COMPLETE
開啟逐條審核

Atomic Claim 407/574 · 2026-02-16_inferencex-v2-nvidia-blackwell-vs::IX2-0426

Claim: 在 64-GPU DP64/EP64 deployment 中,每顆 GPU 都會儲存完全相同的一份約 40B non-expert parameters。
Frame: COMPARISON · Mode: ASSERTED · Mapping: PARTIAL
開啟逐條審核

Adding tensor parallelism within groups of GPUs addresses this. In an EP64/DP8/TP8 configuration, the 64 GPUs are organized into 8 DP groups of 8 GPUs each. Within each TP group, the attention projections, shared expert, normalization, and LM head are sharded 8 ways, so each GPU holds only 1/8th of the non-expert weights. Across the full cluster, the 256 experts are still distributed one-per-4-GPUs as before.

解法是在 GPU group 內加入 tensor parallelism。以 EP64/DP8/TP8 為例,64 顆 GPU 被組成 8 個 DP group,每組 8 GPU。每個 TP group 內,attention projection、shared expert、normalization、LM head 都被 shard 8 份,因此每顆 GPU 只放 1/8 non-expert weight。整個 cluster 的 256 個 expert 則仍和前面一樣分散,平均每 4 顆 GPU 對應一組 expert allocation。

Atomic Claim 408/574 · 2026-02-16_inferencex-v2-nvidia-blackwell-vs::IX2-0427

Claim: 可在 GPUs 子群組內加入 tensor parallelism 解決這個問題。
Frame: ATTRIBUTE · Mode: ASSERTED · Mapping: COMPLETE
開啟逐條審核

Atomic Claim 409/574 · 2026-02-16_inferencex-v2-nvidia-blackwell-vs::IX2-0428

Claim: 在 EP64/DP8/TP8 configuration 中,64 顆 GPUs 被分成 8 個 DP groups,每組 8 顆 GPUs
Frame: ATTRIBUTE · Mode: ASSERTED · Mapping: COMPLETE
開啟逐條審核

Atomic Claim 410/574 · 2026-02-16_inferencex-v2-nvidia-blackwell-vs::IX2-0429

Claim: 在每個 TP group 內,attention projections、shared expert 與 normalization 都會分片。
Frame: ATTRIBUTE · Mode: ASSERTED · Mapping: COMPLETE
開啟逐條審核

Atomic Claim 411/574 · 2026-02-16_inferencex-v2-nvidia-blackwell-vs::IX2-0430

Claim: LM head 也會被分成 8 份,因此每顆 GPU 只需持有 1/8 的 non-expert weights。
Frame: ATTRIBUTE · Mode: ASSERTED · Mapping: COMPLETE
開啟逐條審核

Atomic Claim 412/574 · 2026-02-16_inferencex-v2-nvidia-blackwell-vs::IX2-0431

Claim: 在整個 cluster 中,256 個 experts 仍維持先前分布方式,也就是每 4 顆 GPUs 對應一組 expert 分配。
Frame: ATTRIBUTE · Mode: ASSERTED · Mapping: COMPLETE
開啟逐條審核

Pure DEP has a single communication pattern: all-to-all for expert routing. Adding TP introduces a second all-reduce within each TP group for the attention and non-expert computations. The key design principle is to place TP groups within a single node, where NVLink or MNNVL provides high-bandwidth interconnect, and run EP/DP across nodes, where the all-to-all communication pattern can tolerate higher latency.

Pure DEP 只有一種主要 communication pattern:expert routing 用 all-to-all。加入 TP 後,attention 與 non-expert computation 會多一個 TP group 內的 all-reduce。關鍵 design principle 是把 TP group 放在單一 node 內,利用 NVLink 或 MNNVL 的 high-bandwidth interconnect;EP/DP 則跨 node 展開,因為 all-to-all pattern 對較高 latency 的容忍度更高。

Atomic Claim 413/574 · 2026-02-16_inferencex-v2-nvidia-blackwell-vs::IX2-0432

Claim: 關於 EP:Pure DEP 只有一種 communication pattern:用於 expert routing 的 all-to-all。
Frame: ATTRIBUTE · Mode: ASSERTED · Mapping: PARTIAL
開啟逐條審核

Atomic Claim 414/574 · 2026-02-16_inferencex-v2-nvidia-blackwell-vs::IX2-0433

Claim: 加入 TP 後,attention 與 non-expert computations 會在每個 TP group 內多出第二種 all-reduce 通訊。
Frame: ATTRIBUTE · Mode: ASSERTED · Mapping: COMPLETE
開啟逐條審核

Atomic Claim 415/574 · 2026-02-16_inferencex-v2-nvidia-blackwell-vs::IX2-0434

Claim: 核心設計原則是把 TP groups 放在單一 node 內,利用 NVLink 或 MNNVL 提供高頻寬 interconnect。
Frame: ATTRIBUTE · Mode: ASSERTED · Mapping: COMPLETE
開啟逐條審核

Atomic Claim 416/574 · 2026-02-16_inferencex-v2-nvidia-blackwell-vs::IX2-0435

Claim: EP/DP 則跨 nodes 執行,因為 all-to-all communication pattern 對較高 latency 的容忍度更高。
Frame: COMPARISON · Mode: ASSERTED · Mapping: PARTIAL
開啟逐條審核

As always, the tradeoff is that of throughput versus latency. TP=8 within a group means those 8 GPUs now share a batch and must synchronize every decode step, reducing effective DP degree from 64 to 8. Per-GPU batching independence on the attention side is lost. But each DP group now processes attention 8x faster per step, since the matmul is split 8 ways across the TP group. Per-token latency drops while peak concurrency also drops, sliding the configuration along the latency-throughput Pareto frontier relative to pure DEP.

一如既往,這仍是 throughput vs latency 的 trade-off。Group 內 TP=8 代表這 8 顆 GPU 現在共享一個 batch,而且每個 decode step 都要同步,effective DP degree 從 64 降到 8;attention side 也失去 per-GPU batching independence。但每個 DP group 的 attention 每 step 會快 8x,因為 matmul 被切成 8 份在 TP group 平行執行。結果是 per-token latency 下降、peak concurrency 也下降,configuration 相較 pure DEP 會沿 latency-throughput Pareto frontier 往另一個位置移動。

Atomic Claim 417/574 · 2026-02-16_inferencex-v2-nvidia-blackwell-vs::IX2-0436

Claim: 關於 EP:這再次回到 throughput 與 latency 之間的 trade-off。
Frame: ATTRIBUTE · Mode: ASSERTED · Mapping: COMPLETE
開啟逐條審核

Atomic Claim 418/574 · 2026-02-16_inferencex-v2-nvidia-blackwell-vs::IX2-0437

Claim: 在一個 group 內使用 TP=8,代表 8 顆 GPUs 必須共用同一 batch,且每個 decode step 都要同步,使有效 DP degree 從 64 降到 8。
Frame: RELATION · Mode: ASSERTED · Mapping: COMPLETE
開啟逐條審核

Atomic Claim 419/574 · 2026-02-16_inferencex-v2-nvidia-blackwell-vs::IX2-0438

Claim: 因此 attention 端每顆 GPU 獨立 batching 的能力會消失。
Frame: ATTRIBUTE · Mode: ASSERTED · Mapping: COMPLETE
開啟逐條審核

Atomic Claim 420/574 · 2026-02-16_inferencex-v2-nvidia-blackwell-vs::IX2-0439

Claim: 但每個 DP group 的 attention step 會快約 8 倍,因為 matmul 被分散到 TP group 中 8 顆 GPU
Frame: COMPARISON · Mode: ASSERTED · Mapping: COMPLETE
開啟逐條審核

Atomic Claim 421/574 · 2026-02-16_inferencex-v2-nvidia-blackwell-vs::IX2-0440

Claim: 關於 EP:因此 per-token latency 下降,但 peak concurrency 也下降,configuration 會沿著 latency-throughput Pareto frontier 相對 pure DEP 移動。
Frame: ATTRIBUTE · Mode: ASSERTED · Mapping: COMPLETE
開啟逐條審核

Disaggregated Prefill

Disaggregated prefill, sometimes referred to as prefill-decode (PD) disaggregation, is the process of performing prefill and decode phases of LLM inference on separate nodes. Prefill occurs when a request is first processed, and a forward pass is computed on all tokens at once, thereby “prefilling” the KV cache for this request. This is a compute-intensive operation as all tokens feed through the forward pass in parallel. Tokens are then generated or “decoded” one at a time, loading the KV cache from HBM at each decode step. This is a memory-intensive process as the growing KV cache is constantly being loaded.

Disaggregated prefill,也稱 prefill-decode(PD)disaggregation,就是把 LLM inference 的 prefill 與 decode phase 放到不同 node。Prefill 發生在 request 第一次被處理時,所有 token 一次通過 forward pass,藉此把該 request 的 KV cache『prefill』完成;因為所有 token 平行進 forward pass,所以這是 compute-intensive operation。之後 token 一次一個進行 decode,每個 decode step 都要從 HBM 載入不斷變大的 KV cache,因此屬於 memory-intensive process。

Atomic Claim 422/574 · 2026-02-16_inferencex-v2-nvidia-blackwell-vs::IX2-0441

Claim: Disaggregated prefill,有時也稱 prefill-decode(PD)disaggregation,是把 LLM inference 的 prefilldecode 階段放到不同 nodes 執行。
Frame: ATTRIBUTE · Mode: ASSERTED · Mapping: PARTIAL
開啟逐條審核

Atomic Claim 423/574 · 2026-02-16_inferencex-v2-nvidia-blackwell-vs::IX2-0442

Claim: Prefill 發生在 request 第一次被處理時。
Frame: ATTRIBUTE · Mode: ASSERTED · Mapping: PARTIAL
開啟逐條審核

Atomic Claim 424/574 · 2026-02-16_inferencex-v2-nvidia-blackwell-vs::IX2-0443

Claim: 系統會一次對全部 tokens 執行 forward pass,藉此為該 request「prefillKV cache
Frame: ATTRIBUTE · Mode: ASSERTED · Mapping: PARTIAL
開啟逐條審核

Atomic Claim 425/574 · 2026-02-16_inferencex-v2-nvidia-blackwell-vs::IX2-0444

Claim: 關於 disaggregated prefill:因為所有 tokens 都同時通過 forward pass,因此這是 compute-intensive operation。
Frame: ATTRIBUTE · Mode: ASSERTED · Mapping: COMPLETE
開啟逐條審核

Atomic Claim 426/574 · 2026-02-16_inferencex-v2-nvidia-blackwell-vs::IX2-0445

Claim: 之後 tokens 會逐一生成或 decode;每個 decode step 都需從 HBM 載入 KV cache
Frame: ATTRIBUTE · Mode: ASSERTED · Mapping: COMPLETE
開啟逐條審核

Atomic Claim 427/574 · 2026-02-16_inferencex-v2-nvidia-blackwell-vs::IX2-0446

Claim: 由於持續載入不斷成長的 KV cachedecode 是 memory-intensive process。
Frame: ATTRIBUTE · Mode: ASSERTED · Mapping: PARTIAL
開啟逐條審核

In traditional single-node inference, engines interleave prefill and decode on the same GPUs. Incoming prefill requests stall in-flight decode batches, increasing both time-to-first-token and inter-token latency. Chunked prefill mitigates this by breaking long prefills into smaller pieces, but the fundamental resource contention remains. Disaggregated prefill eliminates this entirely!

傳統 single-node inference 會在同一批 GPU 交錯執行 prefill 與 decode。新的 prefill request 一進來,就會 stall 正在進行的 decode batch,同時拉高 time-to-first-token 與 inter-token latency。Chunked prefill 可以把長 prefill 切成較小片段、減輕問題,但根本上的 resource contention 還是在。Disaggregated prefill 則直接把這個問題消掉。

Atomic Claim 428/574 · 2026-02-16_inferencex-v2-nvidia-blackwell-vs::IX2-0447

Claim: 傳統 single-node inference engine 會讓 prefilldecode 交錯執行在同一批 GPUs 上。
Frame: ATTRIBUTE · Mode: ASSERTED · Mapping: PARTIAL
開啟逐條審核

Atomic Claim 429/574 · 2026-02-16_inferencex-v2-nvidia-blackwell-vs::IX2-0448

Claim: 新的 prefill requests 會阻塞進行中的 decode batches,同時拉高 time-to-first-token 與 inter-token latency。
Frame: ATTRIBUTE · Mode: ASSERTED · Mapping: COMPLETE
開啟逐條審核

Atomic Claim 430/574 · 2026-02-16_inferencex-v2-nvidia-blackwell-vs::IX2-0449

Claim: Chunked prefill 可把長 prefill 切成較小 pieces,減輕這個問題。
Frame: COMPARISON · Mode: ASSERTED · Mapping: PARTIAL
開啟逐條審核

Atomic Claim 431/574 · 2026-02-16_inferencex-v2-nvidia-blackwell-vs::IX2-0450

Claim: 關於 Prefill:但底層 resource contention 仍然存在。
Frame: ATTRIBUTE · Mode: ASSERTED · Mapping: PARTIAL
開啟逐條審核

Atomic Claim 432/574 · 2026-02-16_inferencex-v2-nvidia-blackwell-vs::IX2-0451

Claim: Disaggregated prefill 可以把這種 contention 完全消除。
Frame: ATTRIBUTE · Mode: ASSERTED · Mapping: PARTIAL
開啟逐條審核

image

Source: DistServe

Disaggregation also enables independent scaling and optimization of each phase. With separate nodes, each phase can be tuned independently: different parallelism strategies, different batch sizes, and different memory allocation ratios. The ratio of prefill to decode nodes can also be matched to the workload’s input-output length ratio. For instance, prefill-dominated workloads (long input, short output e.g., summarization, RAG, agentic coding with large context windows) allocate more prefill instances. Decode-dominated workloads (short input, long output e.g., chain-of-thought reasoning, long-form generation) allocate more decode instances. Workloads with high cache hit rates also tend toward more decode, since reused KV cache entries from shared system prompts or multi-turn conversation history skip prefill entirely.

Disaggregation 還可以讓兩個 phase 各自獨立 scale 與 optimization。拆成不同 node 後,prefill、decode 可以分別採不同 parallelism strategy、batch size、memory allocation ratio;prefill node 與 decode node 的比例也能依 workload 的 input-output length ratio 調整。例如 prefill-dominated workload——長 input、短 output,如 summarization、RAG、large-context agentic coding——就分配更多 prefill instance;decode-dominated workload——短 input、長 output,如 chain-of-thought reasoning、long-form generation——則放更多 decode instance。Cache hit rate 很高的 workload 也通常偏向 decode,因為 shared system prompt 或 multi-turn conversation history 的 KV cache 可以直接 reuse,完全跳過 prefill。

Atomic Claim 433/574 · 2026-02-16_inferencex-v2-nvidia-blackwell-vs::IX2-0452

Claim: 關於 disaggregated prefill:Disaggregation 也讓兩個階段可以各自獨立 scale 與 optimize。
Frame: ATTRIBUTE · Mode: ASSERTED · Mapping: COMPLETE
開啟逐條審核

Atomic Claim 434/574 · 2026-02-16_inferencex-v2-nvidia-blackwell-vs::IX2-0453

Claim:disaggregated serving 中,prefilldecode 可以採用不同 parallelism strategies。
Frame: NARY_RELATION · Mode: ASSERTED · Mapping: COMPLETE
開啟逐條審核

Atomic Claim 435/574 · 2026-02-16_inferencex-v2-nvidia-blackwell-vs::IX2-0454

Claim:disaggregated serving 中,prefilldecode 可以使用不同 batch sizes。
Frame: NARY_RELATION · Mode: ASSERTED · Mapping: COMPLETE
開啟逐條審核

Atomic Claim 436/574 · 2026-02-16_inferencex-v2-nvidia-blackwell-vs::IX2-0455

Claim:disaggregated serving 中,prefilldecode 可以使用不同 memory-allocation ratios。
Frame: NARY_RELATION · Mode: ASSERTED · Mapping: COMPLETE
開啟逐條審核

Atomic Claim 437/574 · 2026-02-16_inferencex-v2-nvidia-blackwell-vs::IX2-0456

Claim: prefilldecode nodes 的比例,也可以根據 workload 的 input-output length ratio 調整。
Frame: ATTRIBUTE · Mode: ASSERTED · Mapping: PARTIAL
開啟逐條審核

Atomic Claim 438/574 · 2026-02-16_inferencex-v2-nvidia-blackwell-vs::IX2-0457

Claim: 例如 prefill-dominated workloads,也就是長 input、短 output 的 summarization、RAG、具大型 context windows 的 agentic coding,會配置更多 prefill instances。
Frame: RELATION · Mode: ASSERTED · Mapping: COMPLETE
開啟逐條審核

Atomic Claim 439/574 · 2026-02-16_inferencex-v2-nvidia-blackwell-vs::IX2-0458

Claim: Decode-dominated workloads,也就是短 input、長 output 的 chain-of-thought reasoning 或 long-form generation,會配置更多 decode instances。
Frame: ATTRIBUTE · Mode: ASSERTED · Mapping: PARTIAL
開啟逐條審核

Atomic Claim 440/574 · 2026-02-16_inferencex-v2-nvidia-blackwell-vs::IX2-0459

Claim: cache hit rate 高的 workload 也通常偏向配置更多 decode,因為 shared system prompts 或 multi-turn conversation history 可重用 KV cache,直接略過 prefill
Frame: NARY_RELATION · Mode: ASSERTED · Mapping: COMPLETE
開啟逐條審核

The key cost of disaggregation is KV cache transfer. After prefill completes, the full KV cache for that request must be transmitted from the prefill node to the decode node before the first decode token can be generated. For a model like DeepSeek R1 with 61 layers and FP8 KV cache, an 8192-token prefill produces roughly 500MB of KV data that must cross the network, adding directly to TTFT. This transfer is performed over RDMA (typically RoCE or InfiniBand) using zero-copy GPU-to-GPU data movement without CPU involvement. Libraries like NIXL (NVIDIA Inference Transfer Library) abstract the data movement layer behind a unified asynchronous API with pluggable backends for UCX, GPUDirect Storage, and other transports. This decouples the inference engine from any specific transfer protocol and enables disaggregation across heterogeneous hardware where prefill and decode instances may span different device types or interconnects.

Disaggregation 最主要的成本是 KV cache transfer。Prefill 完成後,該 request 的完整 KV cache 必須先從 prefill node 傳到 decode node,才能產生第一個 decode token。以 DeepSeek R1、61 layers、FP8 KV cache 為例,8192-token prefill 大約會產生 500MB KV data,全部都要跨 network 傳輸,直接增加 TTFT。這個 transfer 通常透過 RDMA(RoCE 或 InfiniBand)做 zero-copy GPU-to-GPU data movement,不經 CPU。NIXL(NVIDIA Inference Transfer Library)這類 library 會用 unified asynchronous API 抽象化 data movement layer,並可插拔 UCX、GPUDirect Storage 等 transport backend。這讓 inference engine 不必綁定特定 transfer protocol,也讓 prefill、decode 可以跨 heterogeneous hardware,在不同 device type 或 interconnect 上做 disaggregation。

Atomic Claim 441/574 · 2026-02-16_inferencex-v2-nvidia-blackwell-vs::IX2-0460

Claim: Disaggregation 的主要成本是 KV cache transfer。
Frame: ATTRIBUTE · Mode: ASSERTED · Mapping: COMPLETE
開啟逐條審核

Atomic Claim 442/574 · 2026-02-16_inferencex-v2-nvidia-blackwell-vs::IX2-0461

Claim: prefill 完成後,該 request 的完整 KV cache 必須先從 prefill node 傳到 decode node,之後才能產生第一個 decode token。
Frame: ATTRIBUTE · Mode: ASSERTED · Mapping: COMPLETE
開啟逐條審核

Atomic Claim 443/574 · 2026-02-16_inferencex-v2-nvidia-blackwell-vs::IX2-0462

Claim: 以 61 layers、FP8 KV cacheDeepSeek R1 為例,8192-token prefill 會產生約 500MB KV data,需要跨網路傳輸,因此直接增加 TTFT
Frame: ATTRIBUTE · Mode: ASSERTED · Mapping: COMPLETE
開啟逐條審核

Atomic Claim 444/574 · 2026-02-16_inferencex-v2-nvidia-blackwell-vs::IX2-0463

Claim: 這項傳輸透過 RDMA,通常是 RoCE 或 InfiniBand,以 zero-copy GPU-to-GPU data movement 完成,不需要 CPU 參與。
Frame: ATTRIBUTE · Mode: ASSERTED · Mapping: PARTIAL
開啟逐條審核

Atomic Claim 445/574 · 2026-02-16_inferencex-v2-nvidia-blackwell-vs::IX2-0464

Claim:NIXLNVIDIA Inference Transfer Library)這類 library,透過統一 asynchronous API 抽象化 data-movement layer,並提供 UCXGPUDirect Storage 與其他 transports 的 pluggable backends。
Frame: NARY_RELATION · Mode: ASSERTED · Mapping: PARTIAL
開啟逐條審核

Atomic Claim 446/574 · 2026-02-16_inferencex-v2-nvidia-blackwell-vs::IX2-0465

Claim: 這使 inference engine 不必綁定特定 transfer protocol,也能支援 heterogeneous hardware 的 disaggregation,讓 prefilldecode instances 跨不同 device types 或 interconnects。
Frame: NARY_RELATION · Mode: ASSERTED · Mapping: COMPLETE
開啟逐條審核

image

Source: Github

Optimizing Inference with Wide EP + Disaggregated Serving

Wide EP and disaggregated prefill are separate techniques that are often used together to achieve Pareto optimal performance. In this section, we walk through real results from InferenceX to build intuition for which combinations of parallelism strategy, wide EP, and disaggregated prefill are appropriate at different interactivity levels.

Wide EP 與 disaggregated prefill 是兩套不同 technique,但常搭配使用來取得 Pareto-optimal performance。本節會用 InferenceX 的真實結果,建立直覺:不同 interactivity level 下,該選哪種 parallelism strategy、WideEP 與 disaggregated prefill 組合。

Atomic Claim 447/574 · 2026-02-16_inferencex-v2-nvidia-blackwell-vs::IX2-0466

Claim: Wide EPdisaggregated prefill 是兩種獨立技術,但經常搭配使用 together,以達到 Pareto-optimal performance。
Frame: ATTRIBUTE · Mode: ASSERTED · Mapping: PARTIAL
開啟逐條審核

Atomic Claim 448/574 · 2026-02-16_inferencex-v2-nvidia-blackwell-vs::IX2-0467

Claim: 本節透過真實 InferenceX 結果,建立在不同 interactivity 水準下如何選擇 parallelism strategy 與 wide EP 的直覺。
Frame: ATTRIBUTE · Mode: ASSERTED · Mapping: COMPLETE
開啟逐條審核

Atomic Claim 449/574 · 2026-02-16_inferencex-v2-nvidia-blackwell-vs::IX2-0468

Claim: 並進一步判斷何時適合使用 disaggregated prefill
Frame: ATTRIBUTE · Mode: ASSERTED · Mapping: PARTIAL
開啟逐條審核

It helps to first understand what parallelism strategies fall on what parts of the Pareto frontier for single-node configurations. Take the example of DeepSeek R1 FP4 8k/1k on a single 8-GPU B200 node with TRT-LLM. The optimal strategy shifts as you move along the frontier, driven primarily by batch size and its effect on expert activation density.

先理解 single-node configuration 的各種 parallelism strategy 分別落在 Pareto frontier 哪個區域會比較容易。以單一 8-GPU B200 node、TRT-LLM 跑 DeepSeek R1 FP4 8k/1k 為例,沿 frontier 移動時 optimal strategy 會跟著變,主要驅動因素就是 batch size,以及 batch size 對 expert activation density 的影響。

Atomic Claim 450/574 · 2026-02-16_inferencex-v2-nvidia-blackwell-vs::IX2-0470

Claim: 以單一 8-GPU B200 node 上、使用 TRT-LLM 執行 DeepSeek R1 FP4 8k/1k 為例。
Frame: NARY_RELATION · Mode: ASSERTED · Mapping: COMPLETE
開啟逐條審核

Atomic Claim 451/574 · 2026-02-16_inferencex-v2-nvidia-blackwell-vs::IX2-0471

Claim: 沿著 frontier 移動時,最佳 strategy 會改變,主要由 batch size 以及它對 expert activation density 的影響所決定。
Frame: ATTRIBUTE · Mode: ASSERTED · Mapping: PARTIAL
開啟逐條審核

At the highest interactivity levels (batch 1-16), pure TP outperforms any configuration involving EP. At low batch sizes, only a small fraction of experts activate per step. With EP, these activations are distributed unevenly across GPUs: at batch 4, only 32 of 256 experts fire, and any given GPU has roughly a low double digit percent chance of receiving zero routed tokens in a given layer. TP avoids this by sharding every expert across all GPUs, so all 8 GPUs participate equally in every expert computation regardless of which experts the router selects. We collected expert activation ratio versus batch size data while profiling DeepSeek R1, which confirms that at batch sizes 16 and below, expert activation per layer is very low.

最高 interactivity(batch 1–16)時,pure TP 會贏過任何包含 EP 的 configuration。Small batch 下,每個 step 只啟用少量 expert;使用 EP 時,activation 會不平均地分散在 GPU 間。例如 batch 4 時,256 個 expert 只有 32 個被觸發,任何一顆 GPU 在某一層完全收不到 routed token 的機率大約落在低雙位數百分比。TP 則把每個 expert 都 shard 到所有 GPU,所以不管 router 選中哪些 expert,8 顆 GPU 都會平均參與 expert computation。我們 profiling DeepSeek R1 時收集的 expert activation ratio vs batch size data,也確認 batch <=16 時每層 expert activation 非常低。

Atomic Claim 452/574 · 2026-02-16_inferencex-v2-nvidia-blackwell-vs::IX2-0472

Claim: 在最高 interactivity 區域(batch 1~16),pure TP 會優於任何包含 EP 的 configuration。
Frame: COMPARISON · Mode: ASSERTED · Mapping: PARTIAL
開啟逐條審核

Atomic Claim 453/574 · 2026-02-16_inferencex-v2-nvidia-blackwell-vs::IX2-0473

Claim: 在低 batch size 下,每個 step 只會啟動少部分 experts。
Frame: ATTRIBUTE · Mode: ASSERTED · Mapping: COMPLETE
開啟逐條審核

Atomic Claim 454/574 · 2026-02-16_inferencex-v2-nvidia-blackwell-vs::IX2-0474

Claim: 使用 EP 時,這些 activations 會不均勻分散到 GPUs;例如 batch 4 時,每層 256 個 experts 中只有 32 個會被啟動。
Frame: ATTRIBUTE · Mode: ASSERTED · Mapping: COMPLETE
開啟逐條審核

Atomic Claim 455/574 · 2026-02-16_inferencex-v2-nvidia-blackwell-vs::IX2-0475

Claim: 在這種情況下,任一 GPU 在某一層完全收不到 routed tokens 的機率約為低雙位數百分比。
Frame: ATTRIBUTE · Mode: ASSERTED · Mapping: PARTIAL
開啟逐條審核

Atomic Claim 456/574 · 2026-02-16_inferencex-v2-nvidia-blackwell-vs::IX2-0476

Claim: TP 可避免這個問題,因為每個 expert 都被分片到所有 GPUs,無論 router 選到哪些 experts,8 顆 GPUs 都會平均參與運算。
Frame: ATTRIBUTE · Mode: ASSERTED · Mapping: COMPLETE
開啟逐條審核

Atomic Claim 457/574 · 2026-02-16_inferencex-v2-nvidia-blackwell-vs::IX2-0477

Claim: InferenceX 在 profiling DeepSeek R1 時收集 expert activation ratio vs batch size,結果確認 batch 16 以下每層 expert activation 都很低。
Frame: COMPARISON · Mode: ASSERTED · Mapping: PARTIAL
開啟逐條審核

image

Source: SemiAnalysis

As we move to slightly lower interactivities, batch sizes remain small enough that expert weights are still sharded via TP rather than EP. The crossover occurs around batch 32, where approximately 50-60% of experts activate per layer. At this density, EP’s load imbalance becomes tolerable and its token-routing overhead is cheaper than the per-expert all-reduce required by TP. Configurations in this range use TEP: tensor parallelism for attention (all GPUs collaborate on each attention computation), expert parallelism for MoE layers (experts assigned to specific GPUs with all-to-all routing). In the highest throughput, lowest interactivity region of the frontier, batch sizes are large (128+) and configurations shift to full DEP: attention weights are fully replicated across all GPUs as independent data-parallel ranks, experts are distributed via EP, and batch capacity is maximized at the cost of per-token latency. (128+) and attention weights are fully replicated across all DP ranks, maximizing throughput.

Interactivity 稍微下降時,batch 仍夠小,因此 expert weight 還是透過 TP shard,而不是 EP。Crossover 大約落在 batch 32,此時每層約 50–60% expert 會被啟用;到了這個 density,EP load imbalance 已可接受,而且 token routing overhead 開始低於 TP 對每個 expert 做 all-reduce 的成本。這一段通常使用 TEP:attention 用 tensor parallelism(所有 GPU 一起做 attention computation),MoE layer 用 expert parallelism(expert 分配給特定 GPU,再用 all-to-all routing)。到了 frontier 的最高 throughput、最低 interactivity 區域,batch size 很大(128+),configuration 轉成 full DEP:attention weight 在所有 GPU 完整 replicated、各自成為 independent data-parallel rank,expert 則用 EP 分散;這會最大化 batch capacity 與 throughput,但犧牲 per-token latency。

Atomic Claim 458/574 · 2026-02-16_inferencex-v2-nvidia-blackwell-vs::IX2-0478

Claim: 當 interactivity 稍微下降時,batch size 仍偏小,因此 expert weights 仍較適合用 TP 分片,而非 EP。
Frame: RELATION · Mode: ASSERTED · Mapping: COMPLETE
開啟逐條審核

Atomic Claim 459/574 · 2026-02-16_inferencex-v2-nvidia-blackwell-vs::IX2-0480

Claim: 到這個密度後,EP 的 load imbalance 已可接受,而 token-routing overhead 也低於 TP 對每個 expert 都需執行的 all-reduce 成本。
Frame: COMPARISON · Mode: ASSERTED · Mapping: PARTIAL
開啟逐條審核

Atomic Claim 460/574 · 2026-02-16_inferencex-v2-nvidia-blackwell-vs::IX2-0481

Claim: 此區間 configurations 使用 TEP:attention 採 tensor parallelism,由所有 GPUs 協同完成每次 attention;MoE layers 則採 expert parallelism,把 experts 分配到特定 GPUs 並透過 all-to-all routing。
Frame: NARY_RELATION · Mode: ASSERTED · Mapping: COMPLETE
開啟逐條審核

Atomic Claim 461/574 · 2026-02-16_inferencex-v2-nvidia-blackwell-vs::IX2-0482

Claim: 在 frontier 中 throughput 最高、interactivity 最低的區域,batch size 很大(128+),configuration 會轉成 full DEP:attention weights 在所有 GPUs 上完整複製成獨立 DP ranks,而 experts 則透過 EP 分散。
Frame: ATTRIBUTE · Mode: ASSERTED · Mapping: COMPLETE
開啟逐條審核

image

Source: SemiAnalysis InferenceX

We observe the same general pattern when extending to wide EP with disaggregated prefill. Prefill and decode run with separate parallelism strategies and node counts, both tuned to the workload and target interactivity level. Take an 8k/1k workload (prefill heavy) at the high-throughput, low-interactivity end of the frontier. Prefill is the bottleneck as each request requires a forward pass of 8192 input tokens, which is computationally expensive. Recipes in this region allocate more prefill nodes than decode (4P1D, 7P2D, 4P3D) to sustain high prefill throughput. These prefill nodes run DEP configurations, replicating attention weights across independent data-parallel ranks so that multiple long-context prefills can be processed simultaneously. Decode nodes are fewer but run wide DEP with large batch sizes by the same principle as with single node.

延伸到 wide EP + disaggregated prefill,也看到相同 general pattern。Prefill、decode 使用各自獨立的 parallelism strategy 與 node count,並依 workload、target interactivity tuning。以 prefill-heavy 的 8k/1k workload,在 frontier 的 high-throughput、low-interactivity 端為例:每個 request 都要對 8192 input token 做 forward pass,compute cost 很高,因此 prefill 成為 bottleneck。這一區 recipe 會配置更多 prefill node、較少 decode node,例如 4P1D、7P2D、4P3D,才能維持高 prefill throughput。Prefill node 使用 DEP,複製 attention weight 到 independent DP rank,讓多個 long-context prefill 同時處理;decode node 數較少,但依 single-node 同樣原則,以 large batch 跑 wide DEP。

Atomic Claim 462/574 · 2026-02-16_inferencex-v2-nvidia-blackwell-vs::IX2-0485

Claim: 把範圍擴展到搭配 disaggregated prefillwide EP 時,也可觀察到相同一般模式。
Frame: COMPARISON · Mode: ASSERTED · Mapping: COMPLETE
開啟逐條審核

Atomic Claim 463/574 · 2026-02-16_inferencex-v2-nvidia-blackwell-vs::IX2-0486

Claim: Prefilldecode 可以分別使用不同 parallelism strategies。
Frame: RELATION · Mode: ASSERTED · Mapping: COMPLETE
開啟逐條審核

Atomic Claim 464/574 · 2026-02-16_inferencex-v2-nvidia-blackwell-vs::IX2-0487

Claim: Prefilldecode 也可以配置不同 node counts。
Frame: RELATION · Mode: ASSERTED · Mapping: COMPLETE
開啟逐條審核

Atomic Claim 465/574 · 2026-02-16_inferencex-v2-nvidia-blackwell-vs::IX2-0488

Claim: Prefill/decode 的 parallelism 與 node counts 會依 workload 與目標 interactivity 調整。
Frame: ATTRIBUTE · Mode: ASSERTED · Mapping: COMPLETE
開啟逐條審核

Atomic Claim 466/574 · 2026-02-16_inferencex-v2-nvidia-blackwell-vs::IX2-0489

Claim: 例如 8k/1k、偏 prefill 的 workload,在 frontier 的高 throughput、低 interactivity 端。
Frame: ATTRIBUTE · Mode: ASSERTED · Mapping: COMPLETE
開啟逐條審核

Atomic Claim 467/574 · 2026-02-16_inferencex-v2-nvidia-blackwell-vs::IX2-0490

Claim: 此時 Prefill 是瓶頸,因為每個 request 都需要對 8192 個 input tokens 做一次 forward pass,運算成本很高。
Frame: ATTRIBUTE · Mode: ASSERTED · Mapping: COMPLETE
開啟逐條審核

Atomic Claim 468/574 · 2026-02-16_inferencex-v2-nvidia-blackwell-vs::IX2-0491

Claim: 因此這個區域的 recipes 會配置比 decode 更多的 prefill nodes,例如 4P1D、7P2D、4P3D,以維持高 prefill throughput。
Frame: COMPARISON · Mode: ASSERTED · Mapping: COMPLETE
開啟逐條審核

Atomic Claim 469/574 · 2026-02-16_inferencex-v2-nvidia-blackwell-vs::IX2-0492

Claim: 這些 prefill nodes 採 DEP configuration,把 attention weights 複製到彼此獨立的 data-parallel ranks,以同時處理多個 long-context prefills。
Frame: ATTRIBUTE · Mode: ASSERTED · Mapping: COMPLETE
開啟逐條審核

Atomic Claim 470/574 · 2026-02-16_inferencex-v2-nvidia-blackwell-vs::IX2-0493

Claim: Decode nodes 數量較少,但同樣依 single-node 的原理,以大 batch size 執行 wide DEP。
Frame: COMPARISON · Mode: ASSERTED · Mapping: PARTIAL
開啟逐條審核

On the low interactivity end of the frontier, there are fewer concurrent requests in flight, so a single prefill instance can keep pace with incoming demand. Yet each request still requires 1024 decode steps, and at high interactivity those steps must be fast. Recipes in this region shift to more decode nodes than prefill (1P3D, 1P4D), with each decode instance running TEP at low batch size. Tensor parallelism on attention minimizes per-step latency by sharding the computation across all GPUs in the instance, while expert parallelism handles MoE routing at the moderate batch sizes where EP load balance is sufficient. Multiple small-batch decode instances, rather than fewer large-batch ones, keep per-token latency low while still providing enough concurrent serving capacity.

在 frontier 的另一端,也就是較高 interactivity、較低 concurrency 時,同時進行的 request 少,一個 prefill instance 就足以跟上 incoming demand;但每個 request 仍有 1024 個 decode step,而且高 interactivity 要求每一步都很快。因此 recipe 會改成 decode node 多於 prefill,例如 1P3D、1P4D,每個 decode instance 在 low batch 下跑 TEP。Attention 用 tensor parallelism,把 computation shard 到 instance 內全部 GPU 以降低 per-step latency;MoE routing 則用 expert parallelism,因為 moderate batch 下 EP load balance 已足夠。使用多個 small-batch decode instance,而不是少數 large-batch instance,可以維持低 per-token latency,同時提供足夠 concurrent serving capacity。

Atomic Claim 471/574 · 2026-02-16_inferencex-v2-nvidia-blackwell-vs::IX2-0494

Claim: 在 frontier 的低 interactivity 端,同時進行中的 requests 較少,因此單一 prefill instance 就能跟上 incoming demand。
Frame: COMPARISON · Mode: ASSERTED · Mapping: PARTIAL
開啟逐條審核

Atomic Claim 472/574 · 2026-02-16_inferencex-v2-nvidia-blackwell-vs::IX2-0495

Claim: 但每個 request 仍需要 1024 個 decode steps。
Frame: ATTRIBUTE · Mode: ASSERTED · Mapping: COMPLETE
開啟逐條審核

Atomic Claim 473/574 · 2026-02-16_inferencex-v2-nvidia-blackwell-vs::IX2-0496

Claim: 而在高 interactivity 下,這些 decode steps 必須非常快。
Frame: ATTRIBUTE · Mode: ASSERTED · Mapping: COMPLETE
開啟逐條審核

Atomic Claim 474/574 · 2026-02-16_inferencex-v2-nvidia-blackwell-vs::IX2-0497

Claim: 因此此區域 recipes 會配置比 prefill 更多的 decode nodes,例如 1P3D、1P4D。
Frame: COMPARISON · Mode: ASSERTED · Mapping: COMPLETE
開啟逐條審核

Atomic Claim 475/574 · 2026-02-16_inferencex-v2-nvidia-blackwell-vs::IX2-0498

Claim: 每個 decode instance 在低 batch size 下執行 TEP。
Frame: ATTRIBUTE · Mode: ASSERTED · Mapping: COMPLETE
開啟逐條審核

Atomic Claim 476/574 · 2026-02-16_inferencex-v2-nvidia-blackwell-vs::IX2-0499

Claim: Attention 採 Tensor parallelism,把運算分散到 instance 中所有 GPUs,以降低每 step latency。
Frame: ATTRIBUTE · Mode: ASSERTED · Mapping: COMPLETE
開啟逐條審核

Atomic Claim 477/574 · 2026-02-16_inferencex-v2-nvidia-blackwell-vs::IX2-0500

Claim: expert parallelism 則在中等 batch size、EP load balance 已足夠時負責 MoE routing
Frame: RELATION · Mode: ASSERTED · Mapping: COMPLETE
開啟逐條審核

Atomic Claim 478/574 · 2026-02-16_inferencex-v2-nvidia-blackwell-vs::IX2-0501

Claim: 以多個小 batch 的 decode instance 取代較少的大 batch instance,可以在維持較低每 token 延遲的同時,提供足夠的並行服務能力。
Frame: COMPARISON · Mode: ASSERTED · Mapping: PARTIAL
開啟逐條審核

image

image

image

Source: SemiAnalysis InferenceX

Dive into DeepSeek R1 Single Node Results

On DeepSeek R1 FP8 1k1k, we see that MI355X is competitive with its counterpart B200 on single node scenarios, despite getting mogged on FP4 multi node scenarios. MI355X (SGLang) even beats B200 (SGLang) in throughput performance at lower interactivity levels. Moreover, MI355X (SGLang) beats B200 (TRT and SGLang) in most cases from a perf/TCO perspective.

DeepSeek R1 FP8 1k1k 中,雖然 MI355X 在 FP4 multi-node scenario 被 mogged,但 single-node 其實能和 B200 競爭。低 interactivity 下,MI355X(SGLang)的 throughput 甚至高於 B200(SGLang);若看 perf/TCO,MI355X(SGLang)在多數情況也優於 B200(TRT 與 SGLang)。

Atomic Claim 479/574 · 2026-02-16_inferencex-v2-nvidia-blackwell-vs::IX2-0502

Claim:DeepSeek R1 FP8 1k1k 測試中,MI355X 在單節點情境下可與對應的 B200 競爭,儘管其在 FP4 多節點情境下明顯落後。
Frame: ATTRIBUTE · Mode: ASSERTED · Mapping: COMPLETE
開啟逐條審核

Atomic Claim 480/574 · 2026-02-16_inferencex-v2-nvidia-blackwell-vs::IX2-0503

Claim: 在較低 interactivity 水準下,MI355XSGLang)的吞吐量甚至高於 B200SGLang)。
Frame: COMPARISON · Mode: ASSERTED · Mapping: COMPLETE
開啟逐條審核

Atomic Claim 481/574 · 2026-02-16_inferencex-v2-nvidia-blackwell-vs::IX2-0504

Claim: 從 perf/TCO 角度來看,MI355XSGLang)在多數情況下優於 B200(TRT 與 SGLang)。
Frame: COMPARISON · Mode: ASSERTED · Mapping: COMPLETE
開啟逐條審核

Unfortunately, the year is 2026, and most frontier labs and inference providers are not running FP8 nor single node inference.

可惜現在已經是 2026 年,多數 frontier lab 與 inference provider 早就不是在跑 FP8,也不是跑 single-node inference。

Atomic Claim 482/574 · 2026-02-16_inferencex-v2-nvidia-blackwell-vs::IX2-0505

Claim: 然而,現在已經是 2026 年。
Frame: CLAIM_ONLY · Mode: ASSERTED · Mapping: CLAIM_ONLY
開啟逐條審核

Atomic Claim 483/574 · 2026-02-16_inferencex-v2-nvidia-blackwell-vs::IX2-0506

Claim: 多數 frontier labs 與 inference providers 並未採用 FP8,也不是以單節點方式執行 inference。
Frame: ATTRIBUTE · Mode: ASSERTED · Mapping: COMPLETE
開啟逐條審核

This result goes to show that AMDs chips are great and can be extremely competitive with Nvidia if only they could move faster on the software front. Speed is the moat.

這個結果其實再次證明 AMD chip 本身很強,只要 software front 能跑得更快,就完全有能力和 Nvidia 極度競爭。Speed is the moat。

Atomic Claim 484/574 · 2026-02-16_inferencex-v2-nvidia-blackwell-vs::IX2-0507

Claim: SemiAnalysis 認為,此結果顯示 AMD 的晶片本身相當優秀。
Frame: ATTRIBUTE · Mode: ASSERTED · Mapping: PARTIAL
開啟逐條審核

Atomic Claim 485/574 · 2026-02-16_inferencex-v2-nvidia-blackwell-vs::IX2-0508

Claim: SemiAnalysis 認為,如果 Nvidia 的競爭對手能在軟體面加快進度,這項結果可望具備非常強的競爭力。
Frame: ATTRIBUTE · Mode: HYPOTHETICAL · Mapping: COMPLETE
開啟逐條審核

Atomic Claim 486/574 · 2026-02-16_inferencex-v2-nvidia-blackwell-vs::IX2-0509

Claim: 關於 DeepSeek R1:速度就是護城河。
Frame: ATTRIBUTE · Mode: ASSERTED · Mapping: PARTIAL
開啟逐條審核

image

Source: SemiAnalysis InferenceX

image

Source: SemiAnalysis InferenceMAX

To that end, we see MI355X fall well behind B200 in performance on FP4:

也正因如此,FP4 performance 上 MI355X 就明顯落後 B200:

Atomic Claim 487/574 · 2026-02-16_inferencex-v2-nvidia-blackwell-vs::IX2-0510

Claim:FP4 效能上,MI355X 明顯落後 B200
Frame: ATTRIBUTE · Mode: ASSERTED · Mapping: COMPLETE
開啟逐條審核

image

Source: SemiAnalysis InferenceX

In comparing DeepSeek R1 FP8 perf between H200 (SGLang) and MI325X (SGLang), not much has changed since our initial release of InferenceXv1 last October. The MI325X data was captured on Feb 12th, 2026 with SGLang 0.5.8 whereas the B200 data was captured Jan 23, 2026 with SGLang 0.5.7.

比較 DeepSeek R1 FP8 的 H200(SGLang)與 MI325X(SGLang),從去年 10 月 InferenceXv1 初次發布到現在沒有太大變化。MI325X data 是 2026/2/12 用 SGLang 0.5.8 擷取;文中比較用的另一組 data 則是在 2026/1/23、SGLang 0.5.7 擷取。

Atomic Claim 488/574 · 2026-02-16_inferencex-v2-nvidia-blackwell-vs::IX2-0511

Claim: 比較 DeepSeek R1 FP8H200SGLang)與 MI325XSGLang)上的效能後,自去年 10 月 InferenceXv1 首次發布以來,整體變化不大。
Frame: COMPARISON · Mode: ASSERTED · Mapping: COMPLETE
開啟逐條審核

Atomic Claim 489/574 · 2026-02-16_inferencex-v2-nvidia-blackwell-vs::IX2-0512

Claim: MI325X 資料於 2026 年 2 月 12 日以 SGLang 0.5.8 擷取;B200 資料則於 2026 年 1 月 23 日以 SGLang 0.5.7 擷取。
Frame: ATTRIBUTE · Mode: ASSERTED · Mapping: PARTIAL
開啟逐條審核

One thing we note is the considerably smaller interactivity range for MI325X than H200, with H200 ranging from 30-90 tok/sec/user whereas MI325X ranges from only 13-35 tok/sec/user. This is problematic for providers who would like to serve users at a broader range of interactivity.

一個值得注意的地方,是 MI325X 的 interactivity range 明顯比 H200 小:H200 約 30–90 tok/sec/user,MI325X 只有約 13–35 tok/sec/user。對希望能在更廣 interactivity 範圍服務使用者的 provider 而言,這是實際問題。

Atomic Claim 490/574 · 2026-02-16_inferencex-v2-nvidia-blackwell-vs::IX2-0513

Claim: MI325X 的 interactivity 範圍明顯小於 H200H200 約為 30–90 tok/sec/user,而 MI325X 僅約 13–35 tok/sec/user。
Frame: ATTRIBUTE · Mode: ASSERTED · Mapping: COMPLETE
開啟逐條審核

Atomic Claim 491/574 · 2026-02-16_inferencex-v2-nvidia-blackwell-vs::IX2-0514

Claim: 關於 DeepSeek R1:SemiAnalysis 認為,這對希望服務更廣 interactivity 範圍使用者的供應商而言是一項問題。
Frame: ATTRIBUTE · Mode: ASSERTED · Mapping: COMPLETE
開啟逐條審核

image

Source: SemiAnalysis InferenceX

GPT-OSS 120B Single Node

MI300X, MI325X, H200, and H100 group in the lower-left of the throughput vs interactivity plot, indicating broadly similar tradeoffs, with Nvidia generally holding a modest lead. The next step up is MI355X, which delivers roughly more than 2x higher token throughput per GPU at a given interactivity level, relative to that first group. Within MI355X, ATOM shifts the curve toward higher throughput at low interactivity, suggesting it prioritizes peak throughput over per-user responsiveness.

在 throughput vs interactivity plot 左下角,MI300X、MI325X、H200、H100 聚成一群,代表它們的 trade-off 大致類似,而 Nvidia 通常略微領先。再上一級是 MI355X,在相同 interactivity 下,token throughput per GPU 大約比前一群高超過 2x。MI355X 內部比較時,ATOM 會把 curve 往低 interactivity、高 throughput 方向推,顯示它更偏好 peak throughput,而不是 per-user responsiveness。

Atomic Claim 492/574 · 2026-02-16_inferencex-v2-nvidia-blackwell-vs::IX2-0515

Claim: 在 throughput 對 interactivity 圖中,MI300XMI325XH200H100 聚集於左下區域,代表整體 trade-off 類似,而 Nvidia 通常略占優勢。
Frame: ATTRIBUTE · Mode: ASSERTED · Mapping: COMPLETE
開啟逐條審核

Atomic Claim 493/574 · 2026-02-16_inferencex-v2-nvidia-blackwell-vs::IX2-0516

Claim: 再上一個層級是 MI355X;在相同 interactivity 水準下,其每顆 GPU 的 token throughput 約為前述第一組的 2 倍以上。
Frame: COMPARISON · Mode: ASSERTED · Mapping: COMPLETE
開啟逐條審核

Atomic Claim 494/574 · 2026-02-16_inferencex-v2-nvidia-blackwell-vs::IX2-0517

Claim:MI355X 上,ATOM 會讓曲線往低 interactivity、高 throughput 的方向移動,顯示其更偏重峰值吞吐量而非單一使用者回應速度。
Frame: RELATION · Mode: ASSERTED · Mapping: COMPLETE
開啟逐條審核

Above that tier sits NVIDIA’s B200 and GB200, which outperform MI355X across the frontier. While B200 and GB200 share the same Blackwell compute die, GB200 achieves a higher throughput–interactivity curve because the platform and serving stack reduce non-compute bottlenecks at scale (interconnect/topology, CPU-GPU coupling, and runtime scheduling), translating into effective scale-out and less overhead per token.

更上一層則是 NVIDIA B200、GB200,它們在整條 frontier 都優於 MI355X。B200、GB200 使用相同 Blackwell compute die,但 GB200 的 throughput-interactivity curve 更高,因為 platform 與 serving stack 在 scale 下減少了 non-compute bottleneck,包括 interconnect/topology、CPU-GPU coupling、runtime scheduling,因此 scale-out 更有效率、每 token overhead 更低。

Atomic Claim 495/574 · 2026-02-16_inferencex-v2-nvidia-blackwell-vs::IX2-0518

Claim: 更高一層則是 NVIDIAB200GB200,兩者在整條 frontier 上都優於 MI355X
Frame: COMPARISON · Mode: ASSERTED · Mapping: PARTIAL
開啟逐條審核

Atomic Claim 496/574 · 2026-02-16_inferencex-v2-nvidia-blackwell-vs::IX2-0519

Claim: 雖然 B200GB200 採用相同的 Blackwell compute die,但 GB200 的 throughput–interactivity 曲線更高,因其平台與 serving stack 能在大規模部署時降低非運算瓶頸,包括 interconnect/topology 與 CPU-GPU coupling。
Frame: COMPARISON · Mode: ASSERTED · Mapping: COMPLETE
開啟逐條審核

Atomic Claim 497/574 · 2026-02-16_inferencex-v2-nvidia-blackwell-vs::IX2-0520

Claim: 關於 B200:runtime scheduling 的改善可轉化為更有效的 scale-out,並降低每個 token 的 overhead。
Frame: ATTRIBUTE · Mode: ASSERTED · Mapping: COMPLETE
開啟逐條審核

image

Source: SemiAnalysis InferenceX

If we add cost into the equation, MI355x becomes more competitive: beating B200 at high throughputs. However, GB200 still takes the cake for being the cheapest choice.

若把 cost 納入,MI355X 競爭力會更好:high-throughput 區域可以擊敗 B200。不過最便宜的選擇仍然是 GB200。

Atomic Claim 498/574 · 2026-02-16_inferencex-v2-nvidia-blackwell-vs::IX2-0521

Claim: 納入成本後,MI355x 的競爭力提升,在高吞吐量區間可優於 B200
Frame: COMPARISON · Mode: ASSERTED · Mapping: COMPLETE
開啟逐條審核

Atomic Claim 499/574 · 2026-02-16_inferencex-v2-nvidia-blackwell-vs::IX2-0522

Claim: 不過,GB200 仍是成本最低的選擇。
Frame: ATTRIBUTE · Mode: ASSERTED · Mapping: COMPLETE
開啟逐條審核

image

Source: SemiAnalysis InferenceX

Turning again to the comparison between B200 and GB200 NVL72, it is obvious the impact NVL72 has. We discussed the impact of the GB200 NVL72’s larger 72 GPU scale-up world size vs the B200’s 8 GPU scale-up world size earlier in this article. The output token throughput per GPU more than doubles in the ~100 tok/s/user interactivity range, showing the impact of the NVL72’s larger scale up domain.

再回到 B200 vs GB200 NVL72,可以非常直觀看到 NVL72 的影響。前面已討論 GB200 NVL72 的 72-GPU scale-up world size,相較 B200 只有 8-GPU world size 的差別。在約 100 tok/s/user interactivity 區間,GB200 NVL72 的 output token throughput per GPU 超過翻倍,正是 large scale-up domain 帶來的效果。

Atomic Claim 500/574 · 2026-02-16_inferencex-v2-nvidia-blackwell-vs::IX2-0523

Claim: 再次比較 B200GB200 NVL72 時,可以明顯看出 NVL72 架構帶來的影響。
Frame: ATTRIBUTE · Mode: ASSERTED · Mapping: COMPLETE
開啟逐條審核

Atomic Claim 501/574 · 2026-02-16_inferencex-v2-nvidia-blackwell-vs::IX2-0524

Claim: GB200 NVL72 的 scale-up world size 為 72 顆 GPU,相較之下 B200 的 scale-up world size 為 8 顆 GPU
Frame: COMPARISON · Mode: ASSERTED · Mapping: COMPLETE
開啟逐條審核

Atomic Claim 502/574 · 2026-02-16_inferencex-v2-nvidia-blackwell-vs::IX2-0525

Claim: 在約 100 tok/s/user 的 interactivity 區間,NVL72 較大的 scale-up domain 使每顆 GPU 的 output token throughput 提升超過 2 倍。
Frame: COMPARISON · Mode: ASSERTED · Mapping: PARTIAL
開啟逐條審核

image

Source: SemiAnalysis InferenceX

Core InferenceX Repo Updates

We have made a few core architectural changes to the InferenceX repository to make it easier to understand and reproduce benchmarks. Additionally, we have fully subscribed to AI usage to maximize productivity and increase developer velocity.

我們對 InferenceX repository 做了幾項核心 architecture change,讓 benchmark 更容易理解與 reproduce。此外,我們也全面導入 AI 使用,以最大化 productivity、提高 developer velocity。

Atomic Claim 503/574 · 2026-02-16_inferencex-v2-nvidia-blackwell-vs::IX2-0526

Claim: SemiAnalysis 對 InferenceX repository 做了數項核心架構調整,以提升 benchmark 的可理解性與可重現性。
Frame: ATTRIBUTE · Mode: ASSERTED · Mapping: COMPLETE
開啟逐條審核

Atomic Claim 504/574 · 2026-02-16_inferencex-v2-nvidia-blackwell-vs::IX2-0527

Claim: SemiAnalysis 也全面導入 AI 工具,以最大化生產力並提高 developer velocity。
Frame: CLAIM_ONLY · Mode: ASSERTED · Mapping: CLAIM_ONLY
開啟逐條審核

Core Changes Since InferenceXv1

One of the main changes we have made since v1 is the cadence with which we perform sweeps. Previously, we were jestermaxing and performed a full sweep over each configuration nightly. However, as we added more chips, disaggregated prefill, wide EP, and other features, we realized that running every single night was way too time consuming and wasteful. Moreover, it’s just not necessary – benchmarks only really need to be re-run when recipes change or a new software version is released.

相較 v1,一個主要改變是 sweep cadence。以前我們有點 jestermaxxing,每晚都把每種 configuration 做完整 sweep;但加入更多 chip、disaggregated prefill、wide EP 與其他 feature 後,我們發現 nightly full sweep 太耗時也太浪費。而且其實根本沒必要——benchmark 真正需要重跑的時機,是 recipe 改變或新 software version 發布。

Atomic Claim 505/574 · 2026-02-16_inferencex-v2-nvidia-blackwell-vs::IX2-0528

Claim: InferenceX v1 之後的一項主要改變,是調整執行 sweeps 的 cadence。
Frame: ATTRIBUTE · Mode: ASSERTED · Mapping: PARTIAL
開啟逐條審核

Atomic Claim 506/574 · 2026-02-16_inferencex-v2-nvidia-blackwell-vs::IX2-0529

Claim: 關於 InferenceX:過去會每晚對所有 configuration 執行一次完整 sweep。
Frame: ATTRIBUTE · Mode: ASSERTED · Mapping: PARTIAL
開啟逐條審核

Atomic Claim 507/574 · 2026-02-16_inferencex-v2-nvidia-blackwell-vs::IX2-0530

Claim: 隨著更多晶片、disaggregated prefillwide EP 被加入測試範圍,sweep 的複雜度提高。
Frame: ATTRIBUTE · Mode: ASSERTED · Mapping: COMPLETE
開啟逐條審核

Atomic Claim 508/574 · 2026-02-16_inferencex-v2-nvidia-blackwell-vs::IX2-0531

Claim: 關於 disaggregated prefill:隨著其他功能持續加入,每晚執行所有 sweep 變得過度耗時且浪費資源。
Frame: ATTRIBUTE · Mode: ASSERTED · Mapping: PARTIAL
開啟逐條審核

Atomic Claim 509/574 · 2026-02-16_inferencex-v2-nvidia-blackwell-vs::IX2-0532

Claim: 關於 InferenceX:Benchmark 實際上只需要在 recipe 改變或新的軟體版本發布時重新執行。
Frame: ATTRIBUTE · Mode: ASSERTED · Mapping: COMPLETE
開啟逐條審核

We now trigger sweeps based on additions to a changelog at the root of the repo. When a developer makes a performance-impacting change to a given config, they add an entry to the changelog listing the affected config along with a brief description of the change. All configs are defined in a master configuration YAML file , which serves as the stateful representation of every data point to be swept, including core settings like ISL/OSL, EP, TP, DP, MTP, and so on. When a PR containing a changelog addition is merged, a workflow parses the referenced config keys, pulls the corresponding sweep definitions from the master config, and fans them out as individual GitHub Actions jobs. The jobs collect all data points for the full sweep and upload the results as artifacts.

現在 sweep 會由 repo root 的 changelog ↗ 觸發。Developer 對某個 config 做出會影響 performance 的修改時,就在 changelog 新增一筆,列出受影響 config 與簡短 change description。所有 config 都定義在一份 master configuration YAML ↗ 中,作為每個待 sweep data point 的 stateful representation,包含 ISL/OSL、EP、TP、DP、MTP 等核心設定。當包含 changelog addition 的 PR merge 後,workflow 會 parse 被引用的 config key,從 master config 取出對應 sweep definition,再 fan-out 成獨立 GitHub Actions job。Job 完整收集 sweep 的所有 data point,最後把 result 上傳成 artifact。

Atomic Claim 510/574 · 2026-02-16_inferencex-v2-nvidia-blackwell-vs::IX2-0533

Claim: 關於 GitHub:目前 sweep 會依據 repository 根目錄 changelog 的新增項目觸發。
Frame: ATTRIBUTE · Mode: ASSERTED · Mapping: PARTIAL
開啟逐條審核

Atomic Claim 511/574 · 2026-02-16_inferencex-v2-nvidia-blackwell-vs::IX2-0535

Claim: 所有 config 都定義在一份 master configuration YAML 中,其中保存所有待 sweep data point 的狀態,包括 ISL/OSL、EP、TP、DP、MTP 等核心設定。
Frame: ATTRIBUTE · Mode: ASSERTED · Mapping: PARTIAL
開啟逐條審核

Atomic Claim 512/574 · 2026-02-16_inferencex-v2-nvidia-blackwell-vs::IX2-0534

Claim: 關於 InferenceX:當開發者對特定 config 做出會影響效能的修改時,會在 changelog 新增一筆紀錄,列出受影響的 config 與變更簡述。
Frame: ATTRIBUTE · Mode: ASSERTED · Mapping: PARTIAL
開啟逐條審核

Atomic Claim 513/574 · 2026-02-16_inferencex-v2-nvidia-blackwell-vs::IX2-0536

Claim: 當包含 changelog 新增項目的 PR 被 merge 後,workflow 會解析所引用的 config keys,從 master config 取出對應 sweep definitions,並展開成個別的 GitHub Actions jobs。
Frame: ATTRIBUTE · Mode: ASSERTED · Mapping: COMPLETE
開啟逐條審核

Atomic Claim 514/574 · 2026-02-16_inferencex-v2-nvidia-blackwell-vs::IX2-0537

Claim: 關於 InferenceX:這些 jobs 會收集完整 sweep 的所有 data points,並把結果上傳為 artifacts。
Frame: ATTRIBUTE · Mode: ASSERTED · Mapping: PARTIAL
開啟逐條審核

Below is a high-level diagram of how InferenceX launches jobs.

下圖是 InferenceX 如何啟動 job 的 high-level diagram。

Atomic Claim 515/574 · 2026-02-16_inferencex-v2-nvidia-blackwell-vs::IX2-0538

Claim: 文章提供一張 InferenceX 啟動 jobs 流程的高階架構圖。
Frame: ATTRIBUTE · Mode: ASSERTED · Mapping: COMPLETE
開啟逐條審核

image

Klaud Cold AI Usage

Shortly after the release of InferenceX v1, we realized how much developer throughput was being left on the table by not utilizing AI more in our InferenceX development. So, we rolled our sleeves up and decided to embrace Claude Code and begin absorbing intelligence, one token at a time to the point that we are currently spending at a 3 million dollars’ worth of Claude intelligence, apply here to join the mission. We started our enlightenment journey when we realized the GitHub Copilot agent was free – at first we couldn’t believe this feature came at no cost! We soon realized that Copilot is terrible and it became apparent why GitHub was giving it away for free. You probably would have had to _pay us _to keep using it.

InferenceX v1 發布不久後,我們發現開發過程沒有大量使用 AI,等於白白丟掉很多 developer throughput。於是乾脆全面擁抱 Claude Code,一個 token 一個 token 地吸收 intelligence;目前消耗速度已達約 $6,000/day。若你也想幫 KPI 往『年化吸收 $3 million Claude intelligence』前進,可以申請加入 mission ↗。我們的 enlightenment journey 一開始其實是發現 GitHub Copilot agent 免費,當下還不敢相信這種功能居然不用錢;很快我們就理解原因:Copilot 實在太糟了。要我們繼續用,搞不好還得反過來付錢給我們。

Atomic Claim 516/574 · 2026-02-16_inferencex-v2-nvidia-blackwell-vs::IX2-0539

Claim:InferenceX v1 發布不久後,團隊發現未更積極使用 AI,使 InferenceX 開發仍有大量 developer throughput 未被利用。
Frame: ATTRIBUTE · Mode: ASSERTED · Mapping: COMPLETE
開啟逐條審核

Atomic Claim 517/574 · 2026-02-16_inferencex-v2-nvidia-blackwell-vs::IX2-0543

Claim: SemiAnalysis 戲稱,若要繼續使用 Copilot,可能反而需要付費給他們。
Frame: CLAIM_ONLY · Mode: HYPOTHETICAL · Mapping: CLAIM_ONLY
開啟逐條審核

Atomic Claim 518/574 · 2026-02-16_inferencex-v2-nvidia-blackwell-vs::IX2-0540

Claim: 團隊因此全面採用 Claude Code,目前相關支出 run rate 約為每天 6,000 美元。
Frame: ATTRIBUTE · Mode: ASSERTED · Mapping: COMPLETE
開啟逐條審核

Atomic Claim 519/574 · 2026-02-16_inferencex-v2-nvidia-blackwell-vs::IX2-0541

Claim: 團隊最初是因發現 GitHub Copilot agent 免費提供,而開始更深入使用 AI coding tools。
Frame: ATTRIBUTE · Mode: ASSERTED · Mapping: PARTIAL
開啟逐條審核

Atomic Claim 520/574 · 2026-02-16_inferencex-v2-nvidia-blackwell-vs::IX2-0542

Claim: SemiAnalysis 認為 Copilot 的使用體驗很差,並以此解釋 GitHub 為何免費提供該功能。
Frame: ATTRIBUTE · Mode: ASSERTED · Mapping: PARTIAL
開啟逐條審核

We had been using Claude Code locally ever since it was released. But recently, we have integrated Claude Code into InferenceX development, using it for the usual tasks such as reviewing PRs, but we also have given it the ability to perform sweeps on clusters. With the workflows we setup, Claude can manually initiate runs, view the results, and iterate. This has enabled us to deploy quick fixes easily on the go via the GitHub app.

Claude Code 發布後我們一直都有在 local 使用;最近則把它正式整合進 InferenceX development。除了 review PR 這類一般工作,我們還讓 Claude 能直接在 cluster 上執行 sweep。透過已建立的 workflow,Claude 可以手動啟動 run、查看結果、再 iterate。這讓我們人在外面時,也能透過 GitHub app 很快部署 quick fix。

Atomic Claim 521/574 · 2026-02-16_inferencex-v2-nvidia-blackwell-vs::IX2-0544

Claim: 團隊自 Claude Code 發布後便一直在本地端使用它。
Frame: ATTRIBUTE · Mode: ASSERTED · Mapping: PARTIAL
開啟逐條審核

Atomic Claim 522/574 · 2026-02-16_inferencex-v2-nvidia-blackwell-vs::IX2-0545

Claim: 近期團隊已把 Claude Code 整合進 InferenceX 開發流程,並用於 PR review 等一般任務。
Frame: RELATION · Mode: ASSERTED · Mapping: COMPLETE
開啟逐條審核

Atomic Claim 523/574 · 2026-02-16_inferencex-v2-nvidia-blackwell-vs::IX2-0546

Claim: 團隊也讓 Claude Code 能夠在 clusters 上執行 sweeps。
Frame: ATTRIBUTE · Mode: ASSERTED · Mapping: PARTIAL
開啟逐條審核

Atomic Claim 524/574 · 2026-02-16_inferencex-v2-nvidia-blackwell-vs::IX2-0547

Claim: 透過既有 workflows,Claude 可以手動啟動 runs、查看結果並反覆迭代。
Frame: ATTRIBUTE · Mode: ASSERTED · Mapping: COMPLETE
開啟逐條審核

Atomic Claim 525/574 · 2026-02-16_inferencex-v2-nvidia-blackwell-vs::IX2-0548

Claim: 這套流程讓團隊可以透過 GitHub app 隨時快速部署修正。
Frame: ATTRIBUTE · Mode: ASSERTED · Mapping: PARTIAL
開啟逐條審核

Another cool use case is using Claude to find recipes for new vLLM/SGLang images. When a new image is released, recipes sometimes need to be updated to achieve optimal performance (new environment variables, modified engine arguments, etc.) With our Claude Code integration, we simply open an issue and ask Claude to search through all commits in the image changelog to find necessary changes to be added to the recipe. This works quite well, and although it’s not perfect, it often gives a good starting point.

另一個很有用的 use case,是叫 Claude 幫新 vLLM/SGLang image 找 recipe。新 image 發布後,為了達到 optimal performance,recipe 有時需要一起更新,例如新增 environment variable、修改 engine argument 等。現在透過 Claude Code integration,我們只要開 issue,叫 Claude 搜完整個 image changelog 的 commit,找出 recipe 必須調整的內容。它不是完美,但效果相當不錯,通常能提供很好的起點。

Atomic Claim 526/574 · 2026-02-16_inferencex-v2-nvidia-blackwell-vs::IX2-0549

Claim: 另一個用途是使用 Claude 尋找新版 vLLMSGLang images 所需的 recipes。
Frame: NARY_RELATION · Mode: ASSERTED · Mapping: COMPLETE
開啟逐條審核

Atomic Claim 527/574 · 2026-02-16_inferencex-v2-nvidia-blackwell-vs::IX2-0552

Claim: 關於 Claude:這種做法整體運作良好;雖然不完美,但通常能提供不錯的起點。
Frame: ATTRIBUTE · Mode: ASSERTED · Mapping: PARTIAL
開啟逐條審核

Atomic Claim 528/574 · 2026-02-16_inferencex-v2-nvidia-blackwell-vs::IX2-0551

Claim: 透過 Claude Code 整合,團隊只需建立 issue,要求 Claude 搜尋 image changelog 的所有 commits,以找出 recipe 必須加入的變更。
Frame: RELATION · Mode: ASSERTED · Mapping: COMPLETE
開啟逐條審核

GitHub Actions

In the spirit of open source, all runs occur on GitHub Actions, so benchmark results are verifiable, transparent, and reproducible. However, GitHub outages have been a constant obstacle to our goals recently. We have seen more unicorns lately than any other animal ! But maybe it’s time for us to touch some grass.

延續 open-source 精神,所有 run 都在 GitHub Actions 執行,因此 benchmark result 可以驗證、透明、可 reproduce。不過最近 GitHub outage 一直妨礙我們達成目標;這陣子看到的 unicorn 恐怕比其他動物都多 ↗。也許我們真的該去 touch some grass 了。

Atomic Claim 529/574 · 2026-02-16_inferencex-v2-nvidia-blackwell-vs::IX2-0553

Claim: 基於 open-source 原則,所有 runs 都在 GitHub Actions 上執行,因此 benchmark results 可驗證、透明且可重現。
Frame: ATTRIBUTE · Mode: ASSERTED · Mapping: PARTIAL
開啟逐條審核

Atomic Claim 530/574 · 2026-02-16_inferencex-v2-nvidia-blackwell-vs::IX2-0555

Claim: 作者以「最近看到的 unicorn 比任何其他動物都多」來調侃 GitHub 錯誤頁面頻繁出現。
Frame: CLAIM_ONLY · Mode: ASSERTED · Mapping: CLAIM_ONLY
開啟逐條審核

Atomic Claim 531/574 · 2026-02-16_inferencex-v2-nvidia-blackwell-vs::IX2-0554

Claim: 近期 GitHub outages 持續成為團隊推進工作的障礙。
Frame: ATTRIBUTE · Mode: ASSERTED · Mapping: PARTIAL
開啟逐條審核

Atomic Claim 532/574 · 2026-02-16_inferencex-v2-nvidia-blackwell-vs::IX2-0556

Claim: 作者以「也許該去 touch some grass」作為對前述情況的自嘲。
Frame: CLAIM_ONLY · Mode: ASSERTED · Mapping: CLAIM_ONLY
開啟逐條審核

Microsoft/GitHub themselves are aware of this and have stopped updating its status page with aggregate uptime numbers and are down to a single 9: 97.36% over the past 90 days. The problem doesn’t seem to go away if you choose to ignore it…

Microsoft/GitHub 自己也知道這件事,甚至已不再在 status page 更新 aggregate uptime number;過去 90 天只剩一個 9:97.36%。假裝問題不存在,看起來並不會讓它自己消失。

Atomic Claim 533/574 · 2026-02-16_inferencex-v2-nvidia-blackwell-vs::IX2-0558

Claim: 文中指出,MicrosoftGitHub 過去 90 天 uptime 為 97.36%,僅剩「single 9」等級。
Frame: ATTRIBUTE · Mode: ASSERTED · Mapping: PARTIAL
開啟逐條審核

Atomic Claim 534/574 · 2026-02-16_inferencex-v2-nvidia-blackwell-vs::IX2-0557

Claim: MicrosoftGitHub 已停止在 status page 更新 aggregate uptime 數字。
Frame: ATTRIBUTE · Mode: ASSERTED · Mapping: PARTIAL
開啟逐條審核

Atomic Claim 535/574 · 2026-02-16_inferencex-v2-nvidia-blackwell-vs::IX2-0559

Claim: 作者諷刺指出,選擇忽略問題並不會讓問題消失。
Frame: CLAIM_ONLY · Mode: INFERRED · Mapping: CLAIM_ONLY
開啟逐條審核

image

Source: Outages project

image

Source: Outages project

All in all, GitHub Actions is just alright. It provides a painfully average experience for developers. It is certainly not meant for launching thousands of jobs across a fleet of hundreds of GPUs. Nevertheless, we have worked closely with some GitHub Actions engineers since our launch to better meet the needs of InferenceX, and we can confidently say they have been a pleasure to work with. Moreover, one of our direct asks was to implement lazy loading for jobs when clicking on a workflow run and, while it did take them a while, they eventually implemented the feature.

總體來說,GitHub Actions 就是『還可以』。對 developer 而言,是非常 painfully average 的體驗,而且顯然不是拿來在數百顆 GPU fleet 上同時啟動數千個 job 的。不過 InferenceX 發布後,我們一直和幾位 GitHub Actions engineer 密切合作,想辦法讓它更符合需求;可以很確定地說,他們合作起來非常愉快。我們其中一個直接 request,是點進 workflow run 時替 job 實作 lazy loading;雖然等了一陣子,最後他們真的把功能做出來了 ↗。

Atomic Claim 536/574 · 2026-02-16_inferencex-v2-nvidia-blackwell-vs::IX2-0560

Claim: 整體而言,作者對 GitHub Actions 的評價僅為普通。
Frame: ATTRIBUTE · Mode: ASSERTED · Mapping: PARTIAL
開啟逐條審核

Atomic Claim 537/574 · 2026-02-16_inferencex-v2-nvidia-blackwell-vs::IX2-0561

Claim: 作者認為 GitHub Actions 給開發者的體驗非常平庸。
Frame: ATTRIBUTE · Mode: ASSERTED · Mapping: PARTIAL
開啟逐條審核

Atomic Claim 538/574 · 2026-02-16_inferencex-v2-nvidia-blackwell-vs::IX2-0562

Claim: GitHub Actions 並不是為了在數百顆 GPUs 組成的 fleet 上啟動數千個 jobs 而設計。
Frame: ATTRIBUTE · Mode: ASSERTED · Mapping: COMPLETE
開啟逐條審核

Atomic Claim 539/574 · 2026-02-16_inferencex-v2-nvidia-blackwell-vs::IX2-0563

Claim:InferenceX 發布以來,團隊持續與部分 GitHub Actions 工程師密切合作,以更符合 InferenceX 的需求。
Frame: RELATION · Mode: ASSERTED · Mapping: COMPLETE
開啟逐條審核

Atomic Claim 540/574 · 2026-02-16_inferencex-v2-nvidia-blackwell-vs::IX2-0564

Claim: 作者表示,與這些 GitHub Actions 工程師合作的經驗很好。
Frame: CLAIM_ONLY · Mode: ASSERTED · Mapping: CLAIM_ONLY
開啟逐條審核

Atomic Claim 541/574 · 2026-02-16_inferencex-v2-nvidia-blackwell-vs::IX2-0566

Claim: 雖然花了一些時間,GitHub 最終仍實作了團隊要求的功能。
Frame: ATTRIBUTE · Mode: ASSERTED · Mapping: PARTIAL
開啟逐條審核

Future of InferenceX

Since the initial release of InferenceX in early October 2025, we have worked hard to continuously improve InferenceX. After release, we spent some time refactoring the codebase to make it more scalable, such that new models and inference techniques can now be added in a “plug and play” fashion. These changes enabled us to seamlessly integrate PD-disagg benchmarks for H100, H200, B200, B300, GB200, GB300, and MI355X. We also added accuracy evaluations to our default benchmark pipeline to ensure visibility into model performance across all configurations.

InferenceX 在 2025 年 10 月初首次發布後,我們持續投入大量工作改善。發布後先花了一段時間 refactor codebase,讓架構更 scalable,現在新 model、新 inference technique 都可以像 plug-and-play 一樣加入。這些修改讓我們能順利整合 H100、H200、B200、B300、GB200、GB300、MI355X 的 PD-disagg benchmark;default benchmark pipeline 也加入 accuracy evaluation,確保所有 configuration 的 model performance 都能被看見。

Atomic Claim 542/574 · 2026-02-16_inferencex-v2-nvidia-blackwell-vs::IX2-0567

Claim:InferenceX 於 2025 年 10 月初首次發布後,團隊持續改善 InferenceX
Frame: ATTRIBUTE · Mode: ASSERTED · Mapping: PARTIAL
開啟逐條審核

Atomic Claim 543/574 · 2026-02-16_inferencex-v2-nvidia-blackwell-vs::IX2-0568

Claim: 關於 InferenceX:發布後,團隊重構 codebase 以提升 scalability,使新的 models 與 inference techniques 可以用「plug and play」方式加入。
Frame: ATTRIBUTE · Mode: ASSERTED · Mapping: COMPLETE
開啟逐條審核

Atomic Claim 544/574 · 2026-02-16_inferencex-v2-nvidia-blackwell-vs::IX2-0569

Claim: 這些改動讓團隊能順利整合 H100H200B200B300GB200GB300MI355X 的 PD-disagg benchmarks。
Frame: NARY_RELATION · Mode: ASSERTED · Mapping: COMPLETE
開啟逐條審核

Atomic Claim 545/574 · 2026-02-16_inferencex-v2-nvidia-blackwell-vs::IX2-0570

Claim: 關於 InferenceX:團隊也把 accuracy evaluations 加入預設 benchmark pipeline,以確保所有 configurations 的 model performance 都可被觀察。
Frame: ATTRIBUTE · Mode: ASSERTED · Mapping: PARTIAL
開啟逐條審核

Although we have made many improvements since our release, there is still much work to be done to achieve the north star goal of providing the most real-world inference benchmarks possible. To achieve this goal, we plan to benchmark on real datasets, add an agentic coding performance benchmark, include more SOTA inference optimizations, benchmark more models, and so much more.

雖然發布以來已改善很多,但距離 North Star——盡可能提供最貼近 real-world inference 的 benchmark——仍有很多工作。我們計畫加入 real dataset、agentic coding performance benchmark、更多 SOTA inference optimization、更多 model,還有其他大量擴充。

Atomic Claim 546/574 · 2026-02-16_inferencex-v2-nvidia-blackwell-vs::IX2-0571

Claim: 關於 InferenceX:儘管發布後已完成許多改進,但距離提供最貼近真實世界 inference benchmark 的核心目標仍有不少工作。
Frame: ATTRIBUTE · Mode: ASSERTED · Mapping: PARTIAL
開啟逐條審核

Atomic Claim 547/574 · 2026-02-16_inferencex-v2-nvidia-blackwell-vs::IX2-0573

Claim: 為達成此目標,團隊計畫使用真實 datasets、加入 agentic coding performance benchmark,並 benchmark 更多 models。
Frame: ATTRIBUTE · Mode: ASSERTED · Mapping: PARTIAL
開啟逐條審核

Atomic Claim 548/574 · 2026-02-16_inferencex-v2-nvidia-blackwell-vs::IX2-0574

Claim: 為達成此目標,團隊計畫使用真實 datasets、加入 agentic coding performance benchmark,並持續加入更多 benchmark 能力與內容。
Frame: ATTRIBUTE · Mode: ASSERTED · Mapping: PARTIAL
開啟逐條審核

Atomic Claim 549/574 · 2026-02-16_inferencex-v2-nvidia-blackwell-vs::IX2-0572

Claim: 為達成此目標,團隊計畫使用真實 datasets 進行 benchmark、加入 agentic coding performance benchmark,並納入更多 SOTA inference optimizations。
Frame: ATTRIBUTE · Mode: ASSERTED · Mapping: PARTIAL
開啟逐條審核

Migration to Multi Turn Real Multi-Turn Chat and Agentic Coding Datasets

Currently, InferenceX uses completely random tokens as input for benchmarking. We then vary the ISL/OSL uniformly subject to the distribution [ISL*0.8, ISL], similarly for OSL. Because of the random data, we disable prefix caching in all our benchmarks, as the expected value of a prefix cache hit rate on completely random data is 0%. Furthermore, all the random data is single-turn, meaning each conversation contains only one prompt and one response. While this provides a good baseline Pareto frontier, it is not a practical benchmark setup that mimics real-world production inference workloads.

目前 InferenceX benchmark input 完全使用 random token,ISL/OSL 則在 [ISL*0.8, ISL](OSL 同理)範圍內均勻變動。由於 data 完全 random,prefix cache hit rate 的 expected value 是 0%,所以所有 benchmark 都關閉 prefix caching。此外,目前 random data 全是 single-turn,每個 conversation 只有一個 prompt、一個 response。這可以建立很好的 baseline Pareto frontier,但並不是能模擬 production inference workload 的實務 benchmark setup。

Atomic Claim 550/574 · 2026-02-16_inferencex-v2-nvidia-blackwell-vs::IX2-0575

Claim: 目前 InferenceX 使用完全隨機的 tokens 作為 benchmark input。
Frame: ATTRIBUTE · Mode: ASSERTED · Mapping: PARTIAL
開啟逐條審核

Atomic Claim 551/574 · 2026-02-16_inferencex-v2-nvidia-blackwell-vs::IX2-0577

Claim: 由於使用隨機資料,所有 benchmark 都停用 prefix caching,因為完全隨機資料的 prefix cache 預期 hit rate 為 0%。
Frame: ATTRIBUTE · Mode: ASSERTED · Mapping: PARTIAL
開啟逐條審核

In the near term, we will create a basic multi-turn benchmark with a dataset like allenai/WildChat-4.8M , which captures real users’ multi-turn conversations. In addition to enabling prefix caching on all scenarios, we will enable KV cache CPU offloading, as this is what we see being done in production workloads. This will more accurately evaluate the strengths and weaknesses of each chip. For instance, MI355X has 288GB HBM3e versus B200s 192GB. Therefore, we expect MI355X to perform better in a high concurrency multiturn scenarios as more memory can be allocated to the KV cache. On the other hand, in scenarios where the GPU KV cache is stressed and blocks are offloaded to the CPU, we expect the GBs to excel as these chips have 900GB/s bidirectional CPU-GPU bandwidth, compared to 128GB/s / 256GB/s on HGX with PCIe 5.0 and 6.0, respectively. Moreover, currently we see AMD’s software for CPU offloading is poor, which may negatively affect performance in the same scenarios.

近期我們會用 allenai/WildChat-4.8M ↗ 這類真實 multi-turn conversation dataset,建立基本 multi-turn benchmark。除了所有 scenario 開啟 prefix caching,也會啟用 KV cache CPU offloading,因為 production workload 就是這樣做,能更準確評估各 chip 的優缺點。例如 MI355X 有 288GB HBM3E、B200 只有 192GB,因此 high-concurrency multi-turn scenario 下,MI355X 應該會更好,因為能分更多 memory 給 KV cache。反過來,在 GPU KV cache 壓力大、block 必須 offload 到 CPU 的情況,GB platform 應會占優勢:CPU-GPU bidirectional bandwidth 可達 900GB/s,而 PCIe 5.0/6.0 HGX 分別只有約 128GB/s/256GB/s。此外目前 AMD 的 CPU offloading software 仍偏弱,可能在同類 scenario 進一步拖累 performance。

Atomic Claim 552/574 · 2026-02-16_inferencex-v2-nvidia-blackwell-vs::IX2-0581

Claim: 除了在所有情境啟用 prefix caching 外,團隊也計畫啟用 KV cache CPU offloading,因為這是 production workloads 中實際採用的做法。
Frame: NARY_RELATION · Mode: ASSERTED · Mapping: COMPLETE
開啟逐條審核

Atomic Claim 553/574 · 2026-02-16_inferencex-v2-nvidia-blackwell-vs::IX2-0583

Claim: 例如,MI355X 配備 288GB HBM3e,而 B200 為 192GB。
Frame: RELATION · Mode: ASSERTED · Mapping: COMPLETE
開啟逐條審核

Atomic Claim 554/574 · 2026-02-16_inferencex-v2-nvidia-blackwell-vs::IX2-0584

Claim: 因此,SemiAnalysis 預期 MI355X 在高 concurrency 的 multi-turn 情境下表現較佳,因為可分配更多記憶體給 KV cache
Frame: COMPARISON · Mode: EXPECTED · Mapping: PARTIAL
開啟逐條審核

Atomic Claim 555/574 · 2026-02-16_inferencex-v2-nvidia-blackwell-vs::IX2-0588

Claim:GPU KV cache 壓力升高、blocks offload 到 CPU 的同一句原文中,本拆分 Claim 僅保留「respectively」片段,需回到 Evidence/原文搭配上下文解讀。
Frame: ATTRIBUTE · Mode: EXPECTED · Mapping: PARTIAL
開啟逐條審核

Atomic Claim 556/574 · 2026-02-16_inferencex-v2-nvidia-blackwell-vs::IX2-0587

Claim:GPU KV cache 壓力升高、blocks offload 到 CPU 的同一句原文中,本拆分 Claim 僅保留「6.0」片段,需回到 Evidence/原文搭配上下文解讀。
Frame: ATTRIBUTE · Mode: EXPECTED · Mapping: PARTIAL
開啟逐條審核

Atomic Claim 557/574 · 2026-02-16_inferencex-v2-nvidia-blackwell-vs::IX2-0585

Claim: 另一方面,當 GPU KV cache 壓力提高、blocks 被 offload 到 CPU 時,SemiAnalysis 預期 GB 系列晶片會具優勢,因其具備 900GB/s 雙向 CPU-GPU bandwidth。
Frame: ATTRIBUTE · Mode: EXPECTED · Mapping: PARTIAL
開啟逐條審核

Atomic Claim 558/574 · 2026-02-16_inferencex-v2-nvidia-blackwell-vs::IX2-0586

Claim:GPU KV cache 壓力升高、blocks offload 到 CPU 的情境下,SemiAnalysis 預期 GB 系列會更有優勢;相較之下,HGX 搭配 PCIe 5.0 時的頻寬為 128GB/s/256GB/s。
Frame: COMPARISON · Mode: EXPECTED · Mapping: PARTIAL
開啟逐條審核

Atomic Claim 559/574 · 2026-02-16_inferencex-v2-nvidia-blackwell-vs::IX2-0589

Claim: SemiAnalysis 認為目前 AMDCPU offloading 軟體表現不佳,可能拖累相同情境下的效能。
Frame: COMPARISON · Mode: HYPOTHETICAL · Mapping: PARTIAL
開啟逐條審核

The point is: real-world multiturn datasets test more SOTA inference engine features and can capture more nuanced and robust performance data across all chips.

重點是:real-world multi-turn dataset 能測到更多 SOTA inference engine feature,也能捕捉各 chip 更細緻、更 robust 的 performance 差異。

Atomic Claim 560/574 · 2026-02-16_inferencex-v2-nvidia-blackwell-vs::IX2-0591

Claim: 真實世界 multi-turn datasets 能測試更多 SOTA inference engine,並針對所有晶片取得更細緻且更具 robustness 的效能資料。
Frame: ATTRIBUTE · Mode: ASSERTED · Mapping: COMPLETE
開啟逐條審核

Atomic Claim 561/574 · 2026-02-16_inferencex-v2-nvidia-blackwell-vs::IX2-0590

Claim: 真實世界 multi-turn datasets 會測試更多 SOTA inference engine 功能。
Frame: ATTRIBUTE · Mode: ASSERTED · Mapping: COMPLETE
開啟逐條審核

With the rise of Claude Code, Codex, and Kimi, it is becoming increasingly important to benchmark performance in agentic coding scenarios. Like above, these scenarios are multi-turn but also include extremely long context conversations as well as tool use. In the next few months, we plan on creating a benchmark suite that will most accurately capture the performance of open models in these agentic coding scenarios across all chips.

隨 Claude Code、Codex、Kimi 興起,agentic coding scenario 的 performance benchmark 越來越重要。這些 workload 同樣是 multi-turn,但還包含超長 context conversation 與 tool use。未來幾個月,我們計畫建立一套 benchmark suite,盡可能準確捕捉 open model 在所有 chip 上跑 agentic coding 時的實際 performance。

Atomic Claim 562/574 · 2026-02-16_inferencex-v2-nvidia-blackwell-vs::IX2-0592

Claim: 隨著 Claude CodeCodex 等工具興起,agentic coding workload 的重要性增加。
Frame: ATTRIBUTE · Mode: ASSERTED · Mapping: COMPLETE
開啟逐條審核

Atomic Claim 563/574 · 2026-02-16_inferencex-v2-nvidia-blackwell-vs::IX2-0593

Claim: Kimi 等工具的興起,使 agentic coding 情境的 benchmark 日益重要。
Frame: ATTRIBUTE · Mode: ASSERTED · Mapping: COMPLETE
開啟逐條審核

Atomic Claim 564/574 · 2026-02-16_inferencex-v2-nvidia-blackwell-vs::IX2-0595

Claim: 未來幾個月,團隊計畫建立 benchmark suite,以更準確衡量各種晶片在 agentic coding 情境下執行 open models 的效能。
Frame: ATTRIBUTE · Mode: ASSERTED · Mapping: COMPLETE
開啟逐條審核

Adding TPU, Trainium and More Models

Currently, we continuously benchmark DeepSeek R1 and GPT OSS 120B (previously Llama 3.1 70B as well). To keep up with the newest model architectures, we plan on adding DeepSeek V3.2 (w/ DSA), DeepSeek V4 on Day 0, Kimi K2.5, Qwen3, GLM5, and many more over the course of the next few months. We will also eventually add multi-modal models and be using EPD & CFD (invented by TogetherAI) optimization too.

目前我們持續 benchmark DeepSeek R1 與 GPT-OSS 120B(之前也包含 Llama 3.1 70B)。為了跟上最新 model architecture,接下來幾個月會加入 DeepSeek V3.2(含 DSA)、DeepSeek V4 Day 0、Kimi K2.5、Qwen3、GLM5 等更多模型;之後也會加入 multi-modal model,並採用 TogetherAI 發明的 EPD、CFD optimization。

Atomic Claim 565/574 · 2026-02-16_inferencex-v2-nvidia-blackwell-vs::IX2-0596

Claim: 目前團隊持續 benchmark DeepSeek R1 與 GPT OSS 120B;過去也包含 Llama 3.1 70B。
Frame: RELATION · Mode: ASSERTED · Mapping: COMPLETE
開啟逐條審核

Atomic Claim 566/574 · 2026-02-16_inferencex-v2-nvidia-blackwell-vs::IX2-0597

Claim: 為跟上最新 model architectures,團隊計畫在未來數月加入 DeepSeek V3.2(含 DSA)、Day 0 支援 DeepSeek V4、Kimi K2.5、Qwen3、GLM5 等模型。
Frame: ATTRIBUTE · Mode: ASSERTED · Mapping: COMPLETE
開啟逐條審核

In addition to new models, we are actively working on adding both TPU and Trainium.

除了新 model,我們也正積極把 TPU 與 Trainium 加進來。

Atomic Claim 567/574 · 2026-02-16_inferencex-v2-nvidia-blackwell-vs::IX2-0599

Claim: 除了新增 models,團隊也正積極加入 TPUTrainium
Frame: ATTRIBUTE · Mode: ASSERTED · Mapping: PARTIAL
開啟逐條審核

Total Cost of Ownership (NVL72, Blackwell, Blackwell Ultra, MI355, Hopper, MI325, MI300)

Looking at capital costs across comparable generations, Nvidia systems tend to have higher capital cost than AMD systems. This is driven mostly by higher compute tray content which is driven by higher GPU pricing – it is well known from their financials that Nvidia enjoys higher margins on their GPUs than other vendors. As an example, MI300X compute tray content sits at ~170K for H100 SXM, and the gap widens further in later generations. MI355X is at ~264K and B300 to ~$344K. That incremental silicon content flows directly into higher server cost, and ultimately higher all-in cluster capex per server.

比較相同世代的 capital cost,Nvidia system 通常高於 AMD,主因是 compute tray content 較高,而背後又主要來自更高 GPU 價格;從財報也很清楚,Nvidia GPU margin 高於其他 vendor。例如 MI300X compute tray content 約 $138K,H100 SXM 約 $170K;後續世代差距更大:MI355X 約 $197K,B200 約 $264K,B300 則到約 $344K。這些額外 silicon content 會直接推高 server cost,最後變成更高的 all-in cluster CapEx per server。

Atomic Claim 568/574 · 2026-02-16_inferencex-v2-nvidia-blackwell-vs::IX2-0600

Claim: 比較同世代系統的資本成本,Nvidia 系統的 capital cost 通常高於 AMD 系統。
Frame: COMPARISON · Mode: ASSERTED · Mapping: COMPLETE
開啟逐條審核

Atomic Claim 569/574 · 2026-02-16_inferencex-v2-nvidia-blackwell-vs::IX2-0601

Claim: 主要原因是 compute tray content 較高,而這又來自較高的 GPU 價格;從財務資料可知,NvidiaGPUs 毛利率高於其他廠商。
Frame: COMPARISON · Mode: ASSERTED · Mapping: COMPLETE
開啟逐條審核

Atomic Claim 570/574 · 2026-02-16_inferencex-v2-nvidia-blackwell-vs::IX2-0602

Claim: 例如,MI300X compute tray content 約為 13.8 萬美元,而 H100 SXM 約為 17 萬美元,且後續世代差距進一步擴大。
Frame: COMPARISON · Mode: ASSERTED · Mapping: COMPLETE
開啟逐條審核

Atomic Claim 571/574 · 2026-02-16_inferencex-v2-nvidia-blackwell-vs::IX2-0603

Claim: MI355X 約為 19.7 萬美元,B200 約升至 26.4 萬美元,B300 則約為 34.4 萬美元。
Frame: ATTRIBUTE · Mode: ASSERTED · Mapping: PARTIAL
開啟逐條審核

This similar dynamic carries over into the Blackwell generation, where an increase in GPU content drives rising total Server cost, in turn driving higher total upfront cluster capex per server, resulting in higher capital cost of ownership.

Blackwell 世代也延續相同 dynamic:GPU content 增加推高 total server cost,進一步提高 upfront cluster CapEx per server,最後變成更高 capital cost of ownership。

Atomic Claim 572/574 · 2026-02-16_inferencex-v2-nvidia-blackwell-vs::IX2-0605

Claim: Blackwell 世代也呈現類似動態:GPU content 增加推升整體 Server cost,再進一步提高每台 server 的 upfront cluster capex 與 capital cost of ownership。
Frame: ATTRIBUTE · Mode: ASSERTED · Mapping: PARTIAL
開啟逐條審核

image

Source: SemiAnalysis AI TCO Model

Across comparable generations, operating costs per GPU are broadly similar because chip TDP is the dominant driver of TCO for operating costs. This goes up as you move from H100s to GB300s, given chip TDP double, driving up operating costs per hour per GPU

相近世代之間,operating cost per GPU 大致相似,因為 operating TCO 最主要驅動因素是 chip TDP。從 H100 一路到 GB300,chip TDP 大約翻倍,因此每 GPU 每小時 operating cost 也跟著上升。

Atomic Claim 573/574 · 2026-02-16_inferencex-v2-nvidia-blackwell-vs::IX2-0606

Claim: 在可比較的世代中,每顆 GPU 的 operating cost 大致相近,因晶片 TDP 是 operating-cost TCO 的主要驅動因素。
Frame: COMPARISON · Mode: ASSERTED · Mapping: PARTIAL
開啟逐條審核

Atomic Claim 574/574 · 2026-02-16_inferencex-v2-nvidia-blackwell-vs::IX2-0607

Claim:H100 升級到 GB300 時,因晶片 TDP 約翻倍,每顆 GPU 的每小時 operating cost 也隨之上升。
Frame: COMPARISON · Mode: ASSERTED · Mapping: COMPLETE
開啟逐條審核

image

Source: SemiAnalysis AI TCO Model