SA Article Coverage Review · 2026-02-16_inferencex-v2-nvidia-blackwell-vs
Coverage Summary
- Source: 開啟原始 SA 文章
- Atomic Claims:
574 - Source blocks:
410 - Blocks with ≥1 Atomic Claim:
166 - Blocks without Atomic Claim:
244 - Unplaced Claims:
0
Coverage Review
請從頭到尾閱讀下方 SA 全文。Atomic Claim 會依 Evidence 在原文出現的位置 inline 插入。
若該英文段落已有繁中翻譯 cache,翻譯只會作為淡色閱讀輔助顯示;不會進入 source、Claim provenance 或 Graph。
沒有 Claim callout 的段落不一定有問題;若內容重要且應形成知識,請記到 Missing Claim Notes。
Missing Claim Notes
- 若看到重要但沒有 Atomic Claim 的段落,請在這裡記錄:
- Section:
- Evidence:
- 為什麼重要/應該抽成什麼 Claim:
SA Full Text + Translation + Atomic Claims
InferenceX v2: NVIDIA Blackwell Vs AMD vs Hopper - Formerly InferenceMAX
Introduction
InferenceXv2 (formerly InferenceMAX) builds on the foundation established by InferenceMAXv1, our open-source, continuously updated inference benchmark ↗ that has set a new standard for AI inference performance and economics. InferenceMAXv1 moved beyond static, point-in-time benchmarks by running continuous tests across hundreds of chips and popular open-source frameworks. Free dashboard available here. ↗
Atomic Claim 1/574 · 2026-02-16_inferencex-v2-nvidia-blackwell-vs::IX2-0001
Claim: InferenceXv2(前稱 InferenceMAX)建立在 InferenceMAXv1 的基礎上;後者是一套開源、持續更新的推論 benchmark,為 AI inference 的效能與經濟性建立了新的標準。
Frame:ATTRIBUTE· Mode:ASSERTED· Mapping:PARTIAL
開啟逐條審核
Atomic Claim 2/574 · 2026-02-16_inferencex-v2-nvidia-blackwell-vs::IX2-0002
Claim: InferenceMAXv1 不再只做單一時間點的靜態 benchmark,而是持續測試數百款晶片與主流開源框架。
Frame:ATTRIBUTE· Mode:ASSERTED· Mapping:COMPLETE
開啟逐條審核
Our benchmark has been widely reproduced, validated and/or supported by almost every major buyer ↗ of compute from Google Cloud ↗ to Microsoft Azure ↗ to Oracle, OpenAI ↗, and many more.
Atomic Claim 3/574 · 2026-02-16_inferencex-v2-nvidia-blackwell-vs::IX2-0003
InferenceXv2 builds on this foundation. It expands coverage to include large scale DeepSeek MoE disaggregated inference (disagg prefill, or simply “disagg”) with wide expert parallelism (wideEP) optimization to **all 6 NVIDIA western GPU SKUs from the past 4 years **as well as to every single AMD western GPU SKU released in the past 3 years – in total InferenceXv2 utilizes close to 1000 frontier GPUs for a full benchmark run across all SKUs.
Atomic Claim 4/574 · 2026-02-16_inferencex-v2-nvidia-blackwell-vs::IX2-0006
Claim: InferenceXv2 涵蓋過去四年間 NVIDIA 在西方市場推出的全部 6 款 GPU SKUs。
Frame:NARY_RELATION· Mode:ASSERTED· Mapping:COMPLETE
開啟逐條審核
Atomic Claim 5/574 · 2026-02-16_inferencex-v2-nvidia-blackwell-vs::IX2-0004
Claim: InferenceXv2 延續並擴展這套基礎。
Frame:ATTRIBUTE· Mode:ASSERTED· Mapping:PARTIAL
開啟逐條審核
Atomic Claim 6/574 · 2026-02-16_inferencex-v2-nvidia-blackwell-vs::IX2-0007
Claim: InferenceXv2 涵蓋過去三年間 AMD 在西方市場推出的所有 GPU SKUs。
Frame:NARY_RELATION· Mode:ASSERTED· Mapping:COMPLETE
開啟逐條審核
Atomic Claim 7/574 · 2026-02-16_inferencex-v2-nvidia-blackwell-vs::IX2-0005
Claim: InferenceXv2 涵蓋大規模 DeepSeek MoE disaggregated inference,並搭配 wide expert parallelism。
Frame:ATTRIBUTE· Mode:ASSERTED· Mapping:PARTIAL
開啟逐條審核
Atomic Claim 8/574 · 2026-02-16_inferencex-v2-nvidia-blackwell-vs::IX2-0008
Claim: 一次完整的 InferenceXv2 benchmark run 會使用接近 1,000 顆 frontier GPUs。
Frame:RELATION· Mode:ASSERTED· Mapping:COMPLETE
開啟逐條審核
With today’s release, InferenceXv2 is now the first suite to benchmark the Blackwell Ultra GB300 NVL72 and B300 across the whole pareto frontier curve, and it is the first third party benchmark to test disagg+wideEP multi-node FP4 and FP8 MI355X performance. In future iterations of InferenceX, we will continue to focus heavily on disaggregated serving with wide expert parallelism as that is what is deployed in production at Frontier AI Labs like OpenAI, Anthropic, xAI, Google Deepmind, DeepSeek as well as advanced API providers like TogetherAI, Baseten, and Fireworks. In this article, we will also break down the system engineering principles and economics in play around the latest Claude Code Fast mode feature ↗.
Atomic Claim 9/574 · 2026-02-16_inferencex-v2-nvidia-blackwell-vs::IX2-0009
Claim: InferenceXv2 是第一套針對 GB300 NVL72 與 B300 進行完整 Pareto frontier curve benchmark 的測試套件。
Frame:ATTRIBUTE· Mode:ASSERTED· Mapping:COMPLETE
開啟逐條審核
Atomic Claim 10/574 · 2026-02-16_inferencex-v2-nvidia-blackwell-vs::IX2-0010
Claim: InferenceXv2 是第一套第三方 benchmark,可測試 multi-node MI355X 在 FP4 與 FP8 下的 disagg+wideEP 效能。
Frame:ATTRIBUTE· Mode:ASSERTED· Mapping:COMPLETE
開啟逐條審核
Atomic Claim 11/574 · 2026-02-16_inferencex-v2-nvidia-blackwell-vs::IX2-0011
Claim: 未來版本的 InferenceX 將持續高度聚焦搭配 wide expert parallelism 的 disaggregated serving。
Frame:NARY_RELATION· Mode:EXPECTED· Mapping:COMPLETE
開啟逐條審核
Atomic Claim 12/574 · 2026-02-16_inferencex-v2-nvidia-blackwell-vs::IX2-0011::SPLIT02
Claim: OpenAI、Anthropic、xAI、Google Deepmind、DeepSeek 等 Frontier AI Labs,以及 TogetherAI、Baseten、Fireworks 等 API providers,在 production 中部署搭配 wide expert parallelism 的 disaggregated serving。
Frame:NARY_RELATION· Mode:ASSERTED· Mapping:PARTIAL
開啟逐條審核
Our benchmark is completely open-source under Apache 2.0 – this means that we are able to move at the same rapid speed at which the AI software ecosystem is advancing. If you like our work and would like to show us some support, please drop a star on our GitHub ↗! We also provide a free data visualizer at https://inferencex.com ↗ for everyone in the ML community to explore the complete dataset themselves.
Atomic Claim 13/574 · 2026-02-16_inferencex-v2-nvidia-blackwell-vs::IX2-0013
Claim: inferencex 也提供免費 data visualizer,讓 ML 社群可自行探索完整 dataset。
Frame:ATTRIBUTE· Mode:ASSERTED· Mapping:PARTIAL
開啟逐條審核
Atomic Claim 14/574 · 2026-02-16_inferencex-v2-nvidia-blackwell-vs::IX2-0012
Claim: 關於 InferenceX:這套 benchmark 完全以 Apache 2.0 開源,因此能以接近 AI software ecosystem 演進的速度快速更新。
Frame:ATTRIBUTE· Mode:ASSERTED· Mapping:PARTIAL
開啟逐條審核
We will add DeepSeekv4 and other popular Chinese frontier models with day 0 support as over the past 6 months, we now have cleaned up a lot of tech debt and are able to move fast with stable infrastructure ↗. We will also be adding TPUv7 Ironwood and Trainium3 to InferenceX later this year! If you want to contribute to our impactful mission while earning a competitive compensation, consider applying here ↗.
Atomic Claim 15/574 · 2026-02-16_inferencex-v2-nvidia-blackwell-vs::IX2-0014
Claim: InferenceX 計畫對 DeepSeekv4 與其他熱門中國 frontier models 提供 day-0 support。
Frame:ATTRIBUTE· Mode:ASSERTED· Mapping:COMPLETE
開啟逐條審核
Atomic Claim 16/574 · 2026-02-16_inferencex-v2-nvidia-blackwell-vs::IX2-0015
Claim: InferenceX 團隊表示,過去六個月已清理大量 technical debt。
Frame:ATTRIBUTE· Mode:ASSERTED· Mapping:COMPLETE
開啟逐條審核
Atomic Claim 17/574 · 2026-02-16_inferencex-v2-nvidia-blackwell-vs::IX2-0016
Claim: InferenceX 團隊表示,目前基礎設施已足夠穩定,可以更快速迭代。
Frame:ATTRIBUTE· Mode:ASSERTED· Mapping:COMPLETE
開啟逐條審核
Atomic Claim 18/574 · 2026-02-16_inferencex-v2-nvidia-blackwell-vs::IX2-0017
Claim: InferenceX 計畫在 2026 年稍晚加入 TPUv7 Ironwood。
Frame:ATTRIBUTE· Mode:ASSERTED· Mapping:PARTIAL
開啟逐條審核
Atomic Claim 19/574 · 2026-02-16_inferencex-v2-nvidia-blackwell-vs::IX2-0018
Claim: InferenceX 計畫在 2026 年稍晚加入 Trainium3。
Frame:ATTRIBUTE· Mode:ASSERTED· Mapping:PARTIAL
開啟逐條審核

Source: InferenceMAX GitHub ↗
Key Observations and Results to Highlight
We see competitive perf per TCO results on FP8 MI355X disagg+wideEP SGLang on AMD compared to FP8 B200 disagg+wideEP SGLang, but when compared to widely used Dynamo TRTLLM B200 FP8, TRT continues to framemog. This is amazing news that AMD SGLang Disagg prefill+wideEP for FP8 is able to match NVIDIA’s SGLang performance.
Atomic Claim 20/574 · 2026-02-16_inferencex-v2-nvidia-blackwell-vs::IX2-0019
Atomic Claim 21/574 · 2026-02-16_inferencex-v2-nvidia-blackwell-vs::IX2-0020
Atomic Claim 22/574 · 2026-02-16_inferencex-v2-nvidia-blackwell-vs::IX2-0021
We also see that for single node aggregated serving, AMD’s SGLang delivers better perf per TCO than NVIDIA’s SGLang for FP8. It is also great to see that AMD has deprecated their second class fork of vllm to move further upstream and closer to delivering first class experience. ↗ Stay tuned for our “State of AMD” article where we talk about the many areas where AMD’s pace of improvement has been rapid & also the areas where the pace of improvement has been lackluster. We recommend that NVIDIA focus even more on SGLang & vLLM ecosystem in addition their TRTLLM engine. Jensen needs to staff more resources & engineers towards contributing open ecosystems like SGLang & vLLM ↗.
Atomic Claim 23/574 · 2026-02-16_inferencex-v2-nvidia-blackwell-vs::IX2-0022
Atomic Claim 24/574 · 2026-02-16_inferencex-v2-nvidia-blackwell-vs::IX2-0023
Atomic Claim 25/574 · 2026-02-16_inferencex-v2-nvidia-blackwell-vs::IX2-0024
Atomic Claim 26/574 · 2026-02-16_inferencex-v2-nvidia-blackwell-vs::IX2-0025
SemiAnalysis InferenceX is free open source software and reader-supported. To receive new posts and support our work consider becoming a free or paid subscriber.
Subscribed
When it comes to the latest inference techniques that are used by the most prominent frontier large-scale inference services (such as disagg prefill+wideEP+FP4), Nvidia absolutely frame mogs with the B200, B300 and ASU frat leader, rack scale GB200/GB300 NVL72 across both SGLang and TRTLLM. Nvidia GPUs also dominate when it comes to energy efficiency, with much lower all-in provisioned picoJoules of energy per token across all workloads.
Atomic Claim 27/574 · 2026-02-16_inferencex-v2-nvidia-blackwell-vs::IX2-0026
Atomic Claim 28/574 · 2026-02-16_inferencex-v2-nvidia-blackwell-vs::IX2-0027
Turning to AMD, we find that the biggest issue with inference on their systems and using their software is composability ↗. That is, many of AMDs inference optimization implementations work well in isolation, but when combined with other optimizations, the result is not as competitive as one would expect. Specifically, the composability of disagg prefill, wideEP and FP4 inference optimizations needs significant improvement.
Atomic Claim 29/574 · 2026-02-16_inferencex-v2-nvidia-blackwell-vs::IX2-0028
Atomic Claim 30/574 · 2026-02-16_inferencex-v2-nvidia-blackwell-vs::IX2-0029
Atomic Claim 31/574 · 2026-02-16_inferencex-v2-nvidia-blackwell-vs::IX2-0030
Claim: 尤其 disagg prefill、wideEP 與 FP4 inference 等最佳化之間的 composability 仍需要大幅改善。
Frame:ATTRIBUTE· Mode:ASSERTED· Mapping:COMPLETE
開啟逐條審核
While performance is competitive on AMD when enabling just a subset of the SOTA inference optimizations, enabling all three major optimizations that labs use, AMD’s performance is currently not competitive with Nvidia’s. We strongly recommend to AMD that they focus heavily on composability of different inference optimizations. We have been told that AMD will start focusing on software composability of FP4+distributed inferencing across their whole software stack. This will happen after Chinese New Year as most of their disagg prefill+wideEP 10x inference engineers are based in China
Atomic Claim 32/574 · 2026-02-16_inferencex-v2-nvidia-blackwell-vs::IX2-0031
Atomic Claim 33/574 · 2026-02-16_inferencex-v2-nvidia-blackwell-vs::IX2-0032
Atomic Claim 34/574 · 2026-02-16_inferencex-v2-nvidia-blackwell-vs::IX2-0033
Claim: SemiAnalysis 得知,AMD 將開始強化整個 software stack 中 FP4 + distributed inferencing 的 software composability。
Frame:ATTRIBUTE· Mode:RUMORED· Mapping:PARTIAL
開啟逐條審核
Atomic Claim 35/574 · 2026-02-16_inferencex-v2-nvidia-blackwell-vs::IX2-0034
Nvidia’s GB300 NVL72 doesn’t disappoint. It achieves up to 100x on FP8 vs FP4 compared to even a strong H100 disagg+wideEP+MTP baseline and 65x on FP8 vs FP8. On H100 vs GB200 NVL72, we see up to 55x realized performance difference at 75 tok/s/user. Rack scale Blackwell NVL72 is framemogging hopper and makes hopper looks like it is jestermaxxing. As Jensen said at GTC 2025, he is chief revenue destroyer. ↗
Atomic Claim 36/574 · 2026-02-16_inferencex-v2-nvidia-blackwell-vs::IX2-0035
Claim: Nvidia 的 GB300 NVL72 效能表現符合甚至超越期待。
Frame:ATTRIBUTE· Mode:ASSERTED· Mapping:PARTIAL
開啟逐條審核
Atomic Claim 37/574 · 2026-02-16_inferencex-v2-nvidia-blackwell-vs::IX2-0036
Atomic Claim 38/574 · 2026-02-16_inferencex-v2-nvidia-blackwell-vs::IX2-0037
Claim: 比較 H100 與 GB200 NVL72,在 75 tok/s/user 下可觀察到最高 55 倍的實際效能差距。
Frame:COMPARISON· Mode:ASSERTED· Mapping:COMPLETE
開啟逐條審核
Atomic Claim 39/574 · 2026-02-16_inferencex-v2-nvidia-blackwell-vs::IX2-0038
Atomic Claim 40/574 · 2026-02-16_inferencex-v2-nvidia-blackwell-vs::IX2-0039
Claim: 如 Jensen 在 GTC 2025 所說,他把自己稱為「chief revenue destroyer」。
Frame:CLAIM_ONLY· Mode:ATTRIBUTED· Mapping:CLAIM_ONLY
開啟逐條審核
At GTC 2024, Jensen claimed that Blackwell will deliver up to 30x perf on inference compared to H100, Jensen under promised & overdelivered on Blackwell inference performance. This should curtail the instances of analysts cracking “Jensen Math” jokes for some time.
Atomic Claim 41/574 · 2026-02-16_inferencex-v2-nvidia-blackwell-vs::IX2-0040
Atomic Claim 42/574 · 2026-02-16_inferencex-v2-nvidia-blackwell-vs::IX2-0040::SPLIT02
Atomic Claim 43/574 · 2026-02-16_inferencex-v2-nvidia-blackwell-vs::IX2-0041
Claim: 這應該會讓分析師一段時間內少開一些「Jensen Math」的玩笑。
Frame:CLAIM_ONLY· Mode:EXPECTED· Mapping:CLAIM_ONLY
開啟逐條審核

Source: SemiAnalysis InferenceX ↗
Acknowledgments and InferenceX™ (formerly InferenceMAX) Initiative Supporters
We would like to thank Jensen Huang and Ian Buck for supporting this open-source effort by providing access to the latest GB300 NVL72 systems along with access to servers representing all GPU SKUs that they have produced for the past four years. We would like to thank the Nvidia team for allowing us to conduct independent benchmarks across this close to 1000 GPUs. Thank you to Jatin Gangani, Kedar Potdar, Sridhar Ramaswamy, Ishan Dhanani, Sahithi Chigurupati, along with many other Nvidia inference engineers for helping to validate and optimize Blackwell & Hopper configurations.
We’re also grateful to Lisa Su and Anush Elangovan for their support of InferenceMAX and for supporting our work with the dozens of AMD engineers like Chun, Andy, Bill, Ramine, Theresa, Parth, etc that contributed to InferenceMAX & upstream vLLM/SGLang bug fixes, as well as for their responsiveness on helping debug and triage AMD exclusive bugs so as to help optimize AMD performance.
We also want to recognize the SGLang, vLLM, and TensorRT-LLM maintainers for building a world-class software stack and open sourcing it to the entire world. You can check their articles on InferenceX here:
SemiAnalysis InferenceMAX: vLLM maintainers & NVIDIA accelerate Blackwell Inference ↗
GPT-OSS Performance Optimizations: Pushing Pareto Frontier ↗
SGLang & NVIDIA Accelerating SemiAnalysis InferenceMAX & GB200 Together ↗
The InferenceX initiative is also supported by many major buyers of compute and prominent members of the ML community including those from OpenAI, Microsoft, vLLM, Tri Dao, PyTorch Foundation, Oracle and more. You can find the full list here ↗.
SemiAnalysis InferenceX is free open source software and reader-supported. To receive new posts and support our work, consider becoming a free or paid subscriber.
Subscribed
A Primer on Important Technical Concepts
In this section, we will give a brief primer on technical concepts that may help the reader better interpret results. Some readers may not need this and can skip directly to our analysis of results. We will take a deeper dive into some of these topics after the results analysis.
Atomic Claim 44/574 · 2026-02-16_inferencex-v2-nvidia-blackwell-vs::IX2-0042
Claim: 本節會先簡要介紹一些有助於理解 benchmark 結果的技術概念。
Frame:CLAIM_ONLY· Mode:ASSERTED· Mapping:CLAIM_ONLY
開啟逐條審核
Atomic Claim 45/574 · 2026-02-16_inferencex-v2-nvidia-blackwell-vs::IX2-0044
Claim: 熟悉相關概念的讀者可以直接跳到結果分析。
Frame:CLAIM_ONLY· Mode:ASSERTED· Mapping:CLAIM_ONLY
開啟逐條審核
Atomic Claim 46/574 · 2026-02-16_inferencex-v2-nvidia-blackwell-vs::IX2-0043
Claim: 部分讀者可能不需要這段基礎介紹。
Frame:CLAIM_ONLY· Mode:HYPOTHETICAL· Mapping:CLAIM_ONLY
開啟逐條審核
Atomic Claim 47/574 · 2026-02-16_inferencex-v2-nvidia-blackwell-vs::IX2-0045
Claim: 在結果分析之後,文章還會對其中一些主題做更深入說明。
Frame:CLAIM_ONLY· Mode:ASSERTED· Mapping:CLAIM_ONLY
開啟逐條審核
Interactivity vs Throughput Tradeoff
The fundamental tradeoff with LLM inference is throughput versus latency. Interactivity (tok/s/user) describes how fast each user of a system receives tokens – it is the inverse of time per output token (TPOT). Throughput (tok/s) describes how many total tokens a system can crank out across all users. One can achieve higher total throughput by batching requests, but each request will be allocated less FLOPs and thus complete slower. This is analogous to the choice of riding a metro bus vs a race car. The metro bus serves many riders, but also makes frequent stops which takes time, but the cost of the metro bus can be amortized across many passengers. The race car can only carry one or two passengers, but it will make few if any additional stops meaning a faster travel time overall, but it is much more expensive to ride per passenger. The metro bus might make more sense for people heading to the park on a weekend, while the race car might be better for bringing a celebrity to their destination. There is no one size fits all solution.
Atomic Claim 48/574 · 2026-02-16_inferencex-v2-nvidia-blackwell-vs::IX2-0048
Claim: Throughput(tok/s)代表整個系統跨所有使用者每秒總共能產生多少 token。
Frame:CLAIM_ONLY· Mode:ASSERTED· Mapping:CLAIM_ONLY
開啟逐條審核
Atomic Claim 49/574 · 2026-02-16_inferencex-v2-nvidia-blackwell-vs::IX2-0046
Atomic Claim 50/574 · 2026-02-16_inferencex-v2-nvidia-blackwell-vs::IX2-0047
Claim: Interactivity(tok/s/user)代表系統中每位使用者收到 token 的速度,等同 time per output token(TPOT)的倒數。
Frame:ATTRIBUTE· Mode:ASSERTED· Mapping:COMPLETE
開啟逐條審核
Atomic Claim 51/574 · 2026-02-16_inferencex-v2-nvidia-blackwell-vs::IX2-0049
Claim: 透過 batching requests 可以提高整體吞吐量。
Frame:CLAIM_ONLY· Mode:ASSERTED· Mapping:CLAIM_ONLY
開啟逐條審核
Atomic Claim 52/574 · 2026-02-16_inferencex-v2-nvidia-blackwell-vs::IX2-0050
Claim: 但每個 request 分配到的 FLOPs 會變少,因此完成速度會更慢。
Frame:CLAIM_ONLY· Mode:ASSERTED· Mapping:CLAIM_ONLY
開啟逐條審核
Atomic Claim 53/574 · 2026-02-16_inferencex-v2-nvidia-blackwell-vs::IX2-0051
Claim: 這可類比成搭乘大眾巴士與賽車之間的選擇。
Frame:CLAIM_ONLY· Mode:ASSERTED· Mapping:CLAIM_ONLY
開啟逐條審核
Atomic Claim 54/574 · 2026-02-16_inferencex-v2-nvidia-blackwell-vs::IX2-0052
Claim: 大眾巴士一次可以服務許多乘客。
Frame:CLAIM_ONLY· Mode:ASSERTED· Mapping:CLAIM_ONLY
開啟逐條審核
Atomic Claim 55/574 · 2026-02-16_inferencex-v2-nvidia-blackwell-vs::IX2-0053
Claim: 但巴士會頻繁停靠,因此需要更多時間。
Frame:CLAIM_ONLY· Mode:ASSERTED· Mapping:CLAIM_ONLY
開啟逐條審核
Atomic Claim 56/574 · 2026-02-16_inferencex-v2-nvidia-blackwell-vs::IX2-0054
Claim: 巴士的固定成本可以由大量乘客共同攤提。
Frame:CLAIM_ONLY· Mode:ASSERTED· Mapping:CLAIM_ONLY
開啟逐條審核
Atomic Claim 57/574 · 2026-02-16_inferencex-v2-nvidia-blackwell-vs::IX2-0056
Claim: 賽車幾乎不用額外停靠,因此整體移動時間更短。
Frame:CLAIM_ONLY· Mode:ASSERTED· Mapping:CLAIM_ONLY
開啟逐條審核
Atomic Claim 58/574 · 2026-02-16_inferencex-v2-nvidia-blackwell-vs::IX2-0057
Claim: 但以每位乘客計算,賽車的成本明顯更高。
Frame:CLAIM_ONLY· Mode:ASSERTED· Mapping:CLAIM_ONLY
開啟逐條審核
Atomic Claim 59/574 · 2026-02-16_inferencex-v2-nvidia-blackwell-vs::IX2-0055
Claim: 賽車通常只能搭載一到兩名乘客。
Frame:CLAIM_ONLY· Mode:ASSERTED· Mapping:CLAIM_ONLY
開啟逐條審核
Atomic Claim 60/574 · 2026-02-16_inferencex-v2-nvidia-blackwell-vs::IX2-0058
Claim: 週末要去公園的大量乘客,可能更適合搭乘巴士。
Frame:CLAIM_ONLY· Mode:HYPOTHETICAL· Mapping:CLAIM_ONLY
開啟逐條審核
Atomic Claim 61/574 · 2026-02-16_inferencex-v2-nvidia-blackwell-vs::IX2-0059
Claim: 若是要快速把名人送到目的地,賽車可能更合適。
Frame:CLAIM_ONLY· Mode:HYPOTHETICAL· Mapping:CLAIM_ONLY
開啟逐條審核
Atomic Claim 62/574 · 2026-02-16_inferencex-v2-nvidia-blackwell-vs::IX2-0060
Claim: 不存在一種適合所有情境的單一最佳方案。
Frame:CLAIM_ONLY· Mode:ASSERTED· Mapping:CLAIM_ONLY
開啟逐條審核

Source: SemiAnalysis
Most benchmark results we will show in this article are InferenceX is a curve. It is important to analyze throughput at various levels of interactivity/latency instead of just looking at maximum achieved throughput (which normally can only be achieved at a single low interactivity). With inference, there is no one size fits all use case. The level of interactivity and throughput needed depends on the use case. For instance, real-time speech models require extremely low latency so that the end user can maintain a natural “conversation” with the LLM, whereas a basic QA chatbot may allow for higher latency. We leave it up to the reader to look at the curve and apply this principle to identify where their use case falls on the throughput-interactivity curve.
Atomic Claim 63/574 · 2026-02-16_inferencex-v2-nvidia-blackwell-vs::IX2-0061
Claim: 大多數 InferenceX benchmark 結果都以曲線呈現。
Frame:ATTRIBUTE· Mode:ASSERTED· Mapping:PARTIAL
開啟逐條審核
Atomic Claim 64/574 · 2026-02-16_inferencex-v2-nvidia-blackwell-vs::IX2-0062
Claim: 分析時應比較不同 interactivity/latency 水準下的 throughput,而不是只看最大吞吐量,因為最大吞吐量通常只能在單一低 interactivity 點達成。
Frame:CLAIM_ONLY· Mode:ASSERTED· Mapping:CLAIM_ONLY
開啟逐條審核
Atomic Claim 65/574 · 2026-02-16_inferencex-v2-nvidia-blackwell-vs::IX2-0063
Claim: 推論不存在適用所有 use case 的單一最佳配置。
Frame:CLAIM_ONLY· Mode:ASSERTED· Mapping:CLAIM_ONLY
開啟逐條審核
Atomic Claim 66/574 · 2026-02-16_inferencex-v2-nvidia-blackwell-vs::IX2-0064
Claim: 所需的 interactivity 與 throughput 取決於實際 use case。
Frame:CLAIM_ONLY· Mode:ASSERTED· Mapping:CLAIM_ONLY
開啟逐條審核
Atomic Claim 67/574 · 2026-02-16_inferencex-v2-nvidia-blackwell-vs::IX2-0065
Atomic Claim 68/574 · 2026-02-16_inferencex-v2-nvidia-blackwell-vs::IX2-0067
Claim: 使用者應根據 benchmark curve 判斷自己的 use case 落在 throughput-interactivity curve 的哪個位置。
Frame:CLAIM_ONLY· Mode:ASSERTED· Mapping:CLAIM_ONLY
開啟逐條審核
The Cost/Perf per TCO vs Interactivity/End-to-End Latency curve mostly follows the Throughput vs Interactivity/End-to-End Latency Curve: More tokens/hour leads to a lower cost per token as fixed $/hour costs are amortized over more tokens produced.
Atomic Claim 69/574 · 2026-02-16_inferencex-v2-nvidia-blackwell-vs::IX2-0068
Claim: Cost/Perf per TCO vs Interactivity/End-to-End Latency 曲線大致跟隨 Throughput vs Interactivity/End-to-End Latency 曲線:每小時產生的 token 越多,固定的每小時成本就能攤在更多 token 上,因此每 token 成本更低。
Frame:CLAIM_ONLY· Mode:ASSERTED· Mapping:CLAIM_ONLY
開啟逐條審核
Prefill and Decode Phases
Inference contains two main phases: prefill and decode. Prefill occurs during the first forward pass of a request’s lifetime. It is computationally intensive since all tokens in the request are processed in parallel. This phase is responsible for “filling up” the KV cache for a sequence. After prefill, responses are generated (or decoded) one token at a time. Each forward pass loads the entire KV cache for a sequence from HBM, while only performing the computation for a single token, making decode memory (bandwidth) intensive.
Atomic Claim 70/574 · 2026-02-16_inferencex-v2-nvidia-blackwell-vs::IX2-0069
Atomic Claim 71/574 · 2026-02-16_inferencex-v2-nvidia-blackwell-vs::IX2-0070
Atomic Claim 72/574 · 2026-02-16_inferencex-v2-nvidia-blackwell-vs::IX2-0074
Atomic Claim 73/574 · 2026-02-16_inferencex-v2-nvidia-blackwell-vs::IX2-0071
Atomic Claim 74/574 · 2026-02-16_inferencex-v2-nvidia-blackwell-vs::IX2-0072
Atomic Claim 75/574 · 2026-02-16_inferencex-v2-nvidia-blackwell-vs::IX2-0073
Atomic Claim 76/574 · 2026-02-16_inferencex-v2-nvidia-blackwell-vs::IX2-0075
When prefill and decode performed on the same engine, prefill constantly disrupts decode batches leading to worse overall performance.
Atomic Claim 77/574 · 2026-02-16_inferencex-v2-nvidia-blackwell-vs::IX2-0076
Disaggregated Prefill
Disaggregated prefill (aka PD disaggregation or simply “disagg”) is the practice of separating the prefill and decode phases across separate pools of GPUs or clusters. These separate prefill and decode pools can be tuned independently and scaled to match the needs of workloads.
Atomic Claim 78/574 · 2026-02-16_inferencex-v2-nvidia-blackwell-vs::IX2-0077
Claim: Disaggregated prefill(也稱 PD disaggregation 或簡稱 disagg)是把 prefill 與 decode 階段分離到不同的 GPUs pools 或 clusters 的做法。
Frame:ATTRIBUTE· Mode:ASSERTED· Mapping:PARTIAL
開啟逐條審核
Atomic Claim 79/574 · 2026-02-16_inferencex-v2-nvidia-blackwell-vs::IX2-0078
Tensor Parallel, Expert Parallel, Data Parallel (TP, EP, DP)
TP allows for maximize interactivity at small batch sizes, but it must carry out an all-reduce at every layer. EP shards experts, exploiting MoE sparsity, with the drawback being an all-to-all collective (which is more costly than simpler collectives like all-reduce) is carried out for MoE layers and can be imbalanced at small batches. DP replicates the entire model (or just parts of a model, like attention) on multiple groups of GPUs (ranks) and then load balances requests among ranks. It is the simplest to scale, but repeats weight loading which can be wasteful at scale.
Atomic Claim 80/574 · 2026-02-16_inferencex-v2-nvidia-blackwell-vs::IX2-0080
Claim: 但 TP 每一層都必須執行一次 all-reduce。
Frame:ATTRIBUTE· Mode:ASSERTED· Mapping:PARTIAL
開啟逐條審核
Atomic Claim 81/574 · 2026-02-16_inferencex-v2-nvidia-blackwell-vs::IX2-0079
Claim: TP 可在小 batch size 下最大化 interactivity。
Frame:ATTRIBUTE· Mode:ASSERTED· Mapping:COMPLETE
開啟逐條審核
Atomic Claim 82/574 · 2026-02-16_inferencex-v2-nvidia-blackwell-vs::IX2-0081
Atomic Claim 83/574 · 2026-02-16_inferencex-v2-nvidia-blackwell-vs::IX2-0082
Claim: 其缺點是 MoE layers 需要執行 all-to-all collective,而這比 all-reduce 等較簡單 collective 成本更高。
Frame:COMPARISON· Mode:ASSERTED· Mapping:COMPLETE
開啟逐條審核
Atomic Claim 84/574 · 2026-02-16_inferencex-v2-nvidia-blackwell-vs::IX2-0084
Tracking Improvements Over Time
One of the main goals of InferenceX is to visualize performance improvements over time. While new chips are released on an O(yearly) cadence, software releases happen on an O(weekly) cadence. Our goal is to constantly update recipes with the latest and greatest software improvements and benchmark the configurations.
Atomic Claim 85/574 · 2026-02-16_inferencex-v2-nvidia-blackwell-vs::IX2-0087
Claim: InferenceX 的主要目標之一,是視覺化效能隨時間的改善。
Frame:ATTRIBUTE· Mode:ASSERTED· Mapping:PARTIAL
開啟逐條審核
Atomic Claim 86/574 · 2026-02-16_inferencex-v2-nvidia-blackwell-vs::IX2-0088
Claim: 關於 InferenceX:新晶片大約以年度 cadence 推出,但 software releases 大約是每週 cadence。
Frame:ATTRIBUTE· Mode:ASSERTED· Mapping:COMPLETE
開啟逐條審核
Atomic Claim 87/574 · 2026-02-16_inferencex-v2-nvidia-blackwell-vs::IX2-0089
Claim: InferenceX 的目標是持續用最新 software improvements 更新 recipes,並重新 benchmark 各種 configurations。
Frame:ATTRIBUTE· Mode:ASSERTED· Mapping:COMPLETE
開啟逐條審核
DeepSeek R1
The AMD team has significantly improved performance for all configurations of SGLang DeepSeek R1 FP4. For the same interactivity, AMD has almost doubled the amount of throughput in the span of less than 2 months. Moreover, we have pushed AMD to upstream performance enhancing changes from their forked SGLang images into the official SGLang image. From December 2025 to January 2026, AMD’s software was improved up to 2x in performance.
Atomic Claim 88/574 · 2026-02-16_inferencex-v2-nvidia-blackwell-vs::IX2-0090
Atomic Claim 89/574 · 2026-02-16_inferencex-v2-nvidia-blackwell-vs::IX2-0091
Atomic Claim 90/574 · 2026-02-16_inferencex-v2-nvidia-blackwell-vs::IX2-0092
Atomic Claim 91/574 · 2026-02-16_inferencex-v2-nvidia-blackwell-vs::IX2-0093

Source: SemiAnalysis InferenceX ↗
In order to continue becoming closer to an first class experience, AMD needs increase their support of vLLM & SGLang maintainers through compute contributions and code contributions & having more reviewers that work for AMD to speed up the review process of AMD PRs into the upstream.
Atomic Claim 92/574 · 2026-02-16_inferencex-v2-nvidia-blackwell-vs::IX2-0094

Source: SemiAnalysis
On the other hand, Nvidia’s results were more consistent, with minor improvements for B200 SGLang over a similar time period.
Atomic Claim 93/574 · 2026-02-16_inferencex-v2-nvidia-blackwell-vs::IX2-0095

Source: SemiAnalysis InferenceX ↗
Many of the mature SKUs had minimal improvements. For example, H200 TRT single node has not changed in performance in the span of 4 months since October, but this is because Hopper support has been excellent since day 1, and performance has close to peak theoretical for this workload all along, making it hard to deliver incremental performance gains.
Atomic Claim 94/574 · 2026-02-16_inferencex-v2-nvidia-blackwell-vs::IX2-0096
Claim: 關於 DeepSeek R1:許多已成熟的 SKUs 幾乎沒有明顯改善。
Frame:ATTRIBUTE· Mode:ASSERTED· Mapping:COMPLETE
開啟逐條審核
Atomic Claim 95/574 · 2026-02-16_inferencex-v2-nvidia-blackwell-vs::IX2-0097
Atomic Claim 96/574 · 2026-02-16_inferencex-v2-nvidia-blackwell-vs::IX2-0098
Atomic Claim 97/574 · 2026-02-16_inferencex-v2-nvidia-blackwell-vs::IX2-0099

Source: SemiAnalysis InferenceX ↗
MI300X and MI325X have seen some improvements, mainly from the most recent SGLang release. Note that for much of the history of InferenceX, AMD was using “private” ROCm images that were not upstreamed, so runs prior to ~Jan 2026 cannot be compared directly to those that are more recent.
Atomic Claim 98/574 · 2026-02-16_inferencex-v2-nvidia-blackwell-vs::IX2-0100
Atomic Claim 99/574 · 2026-02-16_inferencex-v2-nvidia-blackwell-vs::IX2-0101
Claim: 需注意,在 InferenceX 很長一段歷史期間,AMD 使用的是尚未 upstream 的「private」ROCm images,因此約 2026 年 1 月以前的測試結果不能直接與近期結果比較。
Frame:NARY_RELATION· Mode:ASSERTED· Mapping:PARTIAL
開啟逐條審核

Source: SemiAnalysis InferenceX ↗
GB200 Dynamo TRT-LLM disagg has seen some significant improvements as well, with a 20% increase in max throughput in the span of a little over 1 month. We also see improvements in the middle interactivities, where wide EP is deployed. This is likely due to maturing wide EP kernels on GB200.
Atomic Claim 100/574 · 2026-02-16_inferencex-v2-nvidia-blackwell-vs::IX2-0102
Atomic Claim 101/574 · 2026-02-16_inferencex-v2-nvidia-blackwell-vs::IX2-0103
Atomic Claim 102/574 · 2026-02-16_inferencex-v2-nvidia-blackwell-vs::IX2-0104
Atomic Claim 103/574 · 2026-02-16_inferencex-v2-nvidia-blackwell-vs::IX2-0105

Source: SemiAnalysis InferenceX ↗
B200 SGLang has seen steady and continuous improvement for both FP4 and FP8 scenarios since our initial launch, with throughput per GPU doubling at some interactivity levels since last October.
Atomic Claim 104/574 · 2026-02-16_inferencex-v2-nvidia-blackwell-vs::IX2-0109
Atomic Claim 105/574 · 2026-02-16_inferencex-v2-nvidia-blackwell-vs::IX2-0107
Atomic Claim 106/574 · 2026-02-16_inferencex-v2-nvidia-blackwell-vs::IX2-0106
Atomic Claim 107/574 · 2026-02-16_inferencex-v2-nvidia-blackwell-vs::IX2-0108

Source: SemiAnalysis InferenceX ↗
For MI355X Disaggregated inference serving, AMD recommends using SGLang with MoRI. MoRI is AMD’s MoE dispatch/combine collective and KV Cache transfer library ↗ built from first principles by AMD’s cracked 10x China-based engineering team. Although MoRI needs much more open CI and testing, we are strong supporters of the direction that MoRI is taking. This is because instead of taking AMD’s historical approach, which was to fork NVIDIA’s NCCL into RCCL, MoRI is built from scratch by taking the lessons from RCCL/NCCL and building an entirely new package from first principles. The use of MoRI has also delivered good speedups in the span of more than a month, with throughput per GPU increasing by more than 20% in the 20-45 tok/s/user interactivity range.
Atomic Claim 108/574 · 2026-02-16_inferencex-v2-nvidia-blackwell-vs::IX2-0110
Atomic Claim 109/574 · 2026-02-16_inferencex-v2-nvidia-blackwell-vs::IX2-0111
Atomic Claim 110/574 · 2026-02-16_inferencex-v2-nvidia-blackwell-vs::IX2-0112
Atomic Claim 111/574 · 2026-02-16_inferencex-v2-nvidia-blackwell-vs::IX2-0113
Atomic Claim 112/574 · 2026-02-16_inferencex-v2-nvidia-blackwell-vs::IX2-0114

Source: SemiAnalysis InferenceX ↗
GPT-OSS 120B
For MI300X and MI325X, we have seen marginal improvements across the board. Some AITER optimizations helped MI300X performance across all interactivities, and switching to the upstream vLLM ROCm image led to improvements.
Atomic Claim 113/574 · 2026-02-16_inferencex-v2-nvidia-blackwell-vs::IX2-0115
Atomic Claim 114/574 · 2026-02-16_inferencex-v2-nvidia-blackwell-vs::IX2-0116
SemiAnalysis InferenceX is free open source software and reader-supported. To receive new posts and support our work, consider becoming a free or paid subscriber.
Subscribed

Source: SemiAnalysis InferenceX ↗
In the case of the MI325X, it appears that not all performance enhancements that were present in the downstream ROCm fork image (used during the October 5th, 2025 run) have made it into the official vLLM ROCm image.
Unfortunately, the MI355X literally still uses a fork of the vLLM 0.10.1 build rocm/7.0:rocm7.0_ubuntu_22.04_vllm_0.10.1_instinct_20250927_rc1). We would love to have seen it updated it by now, but unfortunately the current official image (0.15.1, at the time this article was written) is not yet optimized for the MI355X and runs into hard errors. We had also run into hard errors crashes on Mi355 for vLLM 0.14. Word on the street is that vLLM 0.16.0 will finally deliver all the changes needed for better MI355X performance.
Atomic Claim 115/574 · 2026-02-16_inferencex-v2-nvidia-blackwell-vs::IX2-0117
Atomic Claim 116/574 · 2026-02-16_inferencex-v2-nvidia-blackwell-vs::IX2-0118
Atomic Claim 117/574 · 2026-02-16_inferencex-v2-nvidia-blackwell-vs::IX2-0119
Atomic Claim 118/574 · 2026-02-16_inferencex-v2-nvidia-blackwell-vs::IX2-0120
Atomic Claim 119/574 · 2026-02-16_inferencex-v2-nvidia-blackwell-vs::IX2-0121
Atomic Claim 120/574 · 2026-02-16_inferencex-v2-nvidia-blackwell-vs::IX2-0122

Source: SemiAnalysis InferenceX ↗
Turning back to Nvidia’s systems, both Hopper and Blackwell saw a steady performance increase between vLLM 0.11.2 and 0.13.0. Soon, we will update recipes for Nvidia GPUs to use the latest vLLM version and we expect even greater performance gains after making the switch. We also observed a performance bump in the latest 1.2.0 version of TRT-LLM.
Atomic Claim 121/574 · 2026-02-16_inferencex-v2-nvidia-blackwell-vs::IX2-0123
Atomic Claim 122/574 · 2026-02-16_inferencex-v2-nvidia-blackwell-vs::IX2-0124
Atomic Claim 123/574 · 2026-02-16_inferencex-v2-nvidia-blackwell-vs::IX2-0125

Source: SemiAnalysis InferenceX ↗

Source: SemiAnalysis InferenceX ↗
Disaggregated Inference Frameworks
NVIDIA uses Dynamo for its disaggregated inference setup. Dynamo ↗ is an inference framework designed for multi-node distributed inference, featuring techniques such as prefill-decode disaggregation, request routing, and KV cache offloading. It is inference-engine agnostic, allowing us to use SGLang and TRT LLM as backends in our benchmark. For AMD, we use SGLang with two different KV cache transfer frameworks: MoRI and Mooncake. MoRI ↗ is a high-performance communication interface focusing on RDMA and GPU integration, offering applications such as network collective operations and expert parallel kernels. Mooncake, which recently joined the PyTorch ecosystem ↗, supports prefill-decode disaggregation and many fault tolerant multi-node features.
Atomic Claim 124/574 · 2026-02-16_inferencex-v2-nvidia-blackwell-vs::IX2-0126
Claim: NVIDIA 的 disaggregated inference 配置使用 Dynamo。
Frame:RELATION· Mode:ASSERTED· Mapping:COMPLETE
開啟逐條審核
Atomic Claim 125/574 · 2026-02-16_inferencex-v2-nvidia-blackwell-vs::IX2-0130
Claim: MoRI 是聚焦 RDMA 與 GPU 整合的高效能 communication interface,可用於 network collective operations 與 expert parallel kernels 等應用。
Frame:NARY_RELATION· Mode:ASSERTED· Mapping:COMPLETE
開啟逐條審核
Atomic Claim 126/574 · 2026-02-16_inferencex-v2-nvidia-blackwell-vs::IX2-0127
Claim: Dynamo 是為 multi-node distributed inference 設計的 inference framework,支援 prefill-decode disaggregation、request routing 與 KV cache offloading 等技術。
Frame:NARY_RELATION· Mode:ASSERTED· Mapping:COMPLETE
開啟逐條審核
Atomic Claim 127/574 · 2026-02-16_inferencex-v2-nvidia-blackwell-vs::IX2-0131
Atomic Claim 128/574 · 2026-02-16_inferencex-v2-nvidia-blackwell-vs::IX2-0128
Atomic Claim 129/574 · 2026-02-16_inferencex-v2-nvidia-blackwell-vs::IX2-0129
DeepSeek Disagg +WideEP Results Deep Dive
At almost all interactivity levels, disagg outperform aggregated inference (grey lines) in terms of total token throughput per GPU. Multi-node disaggregrated prefill framemogs single node aggregrated serving.
Atomic Claim 130/574 · 2026-02-16_inferencex-v2-nvidia-blackwell-vs::IX2-0132
Atomic Claim 131/574 · 2026-02-16_inferencex-v2-nvidia-blackwell-vs::IX2-0133

Source: SemiAnalysis InferenceX ↗
Nvidia continues to push new updates for B200/GB200 FP8. The latest data on DeepSeek FP8 B200 TRT single node (both MTP enabled/disabled) vs GB200 Dynamo+TRT disagg (both MTP enabled/disabled). This indicates consistent engineering effort to improve rack-scale inference software and wideEP kernels.
Atomic Claim 132/574 · 2026-02-16_inferencex-v2-nvidia-blackwell-vs::IX2-0134
Atomic Claim 133/574 · 2026-02-16_inferencex-v2-nvidia-blackwell-vs::IX2-0135

Source: SemiAnalysis InferenceX ↗
When comparing MI355X disaggregated inference vs aggregated inference, we noticed a similar pattern. Disaggregated inference only overtakes aggregated inference at low interactivity, high batch sizes. This is true across FP4, and it is likely due to poorly optimized kernels.
Atomic Claim 134/574 · 2026-02-16_inferencex-v2-nvidia-blackwell-vs::IX2-0137
Claim: 比較 MI355X disaggregated inference 與 aggregated inference 時,也觀察到類似模式。
Frame:ATTRIBUTE· Mode:ASSERTED· Mapping:PARTIAL
開啟逐條審核
Atomic Claim 135/574 · 2026-02-16_inferencex-v2-nvidia-blackwell-vs::IX2-0138
Claim: Disaggregated inference 只有在低 interactivity、高 batch size 時才會超越 aggregated inference。
Frame:ATTRIBUTE· Mode:ASSERTED· Mapping:COMPLETE
開啟逐條審核
Atomic Claim 136/574 · 2026-02-16_inferencex-v2-nvidia-blackwell-vs::IX2-0139
Atomic Claim 137/574 · 2026-02-16_inferencex-v2-nvidia-blackwell-vs::IX2-0140

Source: SemiAnalysis InferenceX ↗
When composing disagg prefill+wideEP with FP4 on the MI355X, we observe suffers subpar performance.
Atomic Claim 138/574 · 2026-02-16_inferencex-v2-nvidia-blackwell-vs::IX2-0141
Although theoretical modeling shows that disagg inference on MI355Xs should perform way better than single node, disagg actually performs worse for higher interactivity levels due to a lack of kernel and collective optimization in the ROCm software stack when composing multiple SOTA inference optimizations together.
Atomic Claim 139/574 · 2026-02-16_inferencex-v2-nvidia-blackwell-vs::IX2-0142
Claim: 雖然理論模型顯示 MI355X 的 disagg inference 應明顯優於 single-node,但在較高 interactivity 下反而更慢,原因是 ROCm software stack 在同時組合多種 SOTA inference optimizations together 時,缺乏足夠的 kernel 與 collective 最佳化。
Frame:COMPARISON· Mode:ASSERTED· Mapping:COMPLETE
開啟逐條審核

Source: SemiAnalysis InferenceX ↗
Nvidia TensorRT LLM and NVL72
TensorRT LLM already serves billions of tokens per hour globally across providers like TogetherAI and other advanced providers, and it has really allowed the GB200 NVL72 and GB300 NVL72 to shine, delivering more than double the performance at high throughput. MTP boosts these results even further, making use of the chips’ full potential.
Atomic Claim 140/574 · 2026-02-16_inferencex-v2-nvidia-blackwell-vs::IX2-0144
Claim: TensorRT LLM 已充分發揮 GB200 NVL72 與 GB300 NVL72 的能力,在高 throughput 區間提供超過兩倍效能。
Frame:COMPARISON· Mode:ASSERTED· Mapping:COMPLETE
開啟逐條審核
Atomic Claim 141/574 · 2026-02-16_inferencex-v2-nvidia-blackwell-vs::IX2-0143
Atomic Claim 142/574 · 2026-02-16_inferencex-v2-nvidia-blackwell-vs::IX2-0145


Source: SemiAnalysis InferenceX
The benefits delivered from the larger world size of the NVL72 family is also evident if we look at cost graphs. At a fixed interactivity level of 60 tok/s/user, each GB200 NVL GPU produces slightly less than triple the number of tokens/s than each B200 does.
Atomic Claim 143/574 · 2026-02-16_inferencex-v2-nvidia-blackwell-vs::IX2-0147
SemiAnalysis InferenceX is free open source software and reader-supported. To receive new posts and support our work, consider becoming a free or paid subscriber.
Subscribed

Source: SemiAnalysis InferenceX ↗
This gap shrinks as interactivity increases. At 130 tok/s/user, the GB200 NVL72 has nearly no advantage and is even more expensive on a $/Million tokens basis. At low batch sizes, the inference workload shrinks enough to fit within a single HGX node’s NVLink domain (i.e. 8 GPUs), and the GB200 NVL72’s larger scale-out advantage starts to disappear.
Atomic Claim 144/574 · 2026-02-16_inferencex-v2-nvidia-blackwell-vs::IX2-0149
Claim: 在 130 tok/s/user 時,GB200 NVL72 幾乎已沒有優勢。
Frame:ATTRIBUTE· Mode:ASSERTED· Mapping:COMPLETE
開啟逐條審核
Atomic Claim 145/574 · 2026-02-16_inferencex-v2-nvidia-blackwell-vs::IX2-0150
Claim: 在 130 tok/s/user 時,GB200 NVL72 以每百萬 token 成本計算甚至更貴。
Frame:ATTRIBUTE· Mode:ASSERTED· Mapping:COMPLETE
開啟逐條審核
Atomic Claim 146/574 · 2026-02-16_inferencex-v2-nvidia-blackwell-vs::IX2-0151
Atomic Claim 147/574 · 2026-02-16_inferencex-v2-nvidia-blackwell-vs::IX2-0152
Claim: 因此 GB200 NVL72 較大 scale-out domain 的優勢開始消失。
Frame:COMPARISON· Mode:ASSERTED· Mapping:PARTIAL
開啟逐條審核

Source: SemiAnalysis InferenceX ↗
Nvidia versus AMD Disagg Prefill
With today’s release of InferenceXv2, for the first time the ML community is able to see a full Pareto frontier for open-source MI355X distributed inference. We show Pareto curves for the B200 and MI355X with and without enabling MTP.
Atomic Claim 148/574 · 2026-02-16_inferencex-v2-nvidia-blackwell-vs::IX2-0153
Claim: 隨著 InferenceXv2 發布,ML 社群首次能看到開源 MI355X distributed inference 的完整 Pareto frontier。
Frame:RELATION· Mode:ASSERTED· Mapping:COMPLETE
開啟逐條審核
Atomic Claim 149/574 · 2026-02-16_inferencex-v2-nvidia-blackwell-vs::IX2-0154
For FP8 disagg prefill, MI355X (MoRI SGLang) is quite competitive with B200 (Dynamo SGLang). Wide EP is not used for either of these configs as all prefill/decode instances run using EP8 at the most. At both ends of the throughput versus interactivity Pareto frontier, MI355X falls behind the B200 slightly. However, MI355X disagg has a slight advantage for certain levels of interactivity in the middle of the curve. Both the B200 and the MI355X benefit from employing MTP, and we observe the same relative performance improvement for both chips when using MTP.
Atomic Claim 150/574 · 2026-02-16_inferencex-v2-nvidia-blackwell-vs::IX2-0155
Atomic Claim 151/574 · 2026-02-16_inferencex-v2-nvidia-blackwell-vs::IX2-0156
Atomic Claim 152/574 · 2026-02-16_inferencex-v2-nvidia-blackwell-vs::IX2-0157
Atomic Claim 153/574 · 2026-02-16_inferencex-v2-nvidia-blackwell-vs::IX2-0158
Atomic Claim 154/574 · 2026-02-16_inferencex-v2-nvidia-blackwell-vs::IX2-0159

Source: SemiAnalysis InferenceX ↗
However, if we were to only measure output (decode) token throughput, we see that output token throughput is much higher for the B200 than for the MI355X at lower interactivity levels. Note that when looking at output token only throughput for disaggregated inference configurations, we normalize throughout by the number of decode GPUs, not total GPUs. It is possible that different numbers of GPUs are used for output when running inference jobs on the B200 and MI355X, but the bottom line is that whatever configuration decode is run on, B200 gets the decode job done faster.
Atomic Claim 155/574 · 2026-02-16_inferencex-v2-nvidia-blackwell-vs::IX2-0160
Atomic Claim 156/574 · 2026-02-16_inferencex-v2-nvidia-blackwell-vs::IX2-0161
Claim: 分析 disaggregated inference 的 output-token-only throughput 時,InferenceX 是以 decode GPUs 數量做正規化,而非全部 GPUs 數量。
Frame:ATTRIBUTE· Mode:ASSERTED· Mapping:PARTIAL
開啟逐條審核
Atomic Claim 157/574 · 2026-02-16_inferencex-v2-nvidia-blackwell-vs::IX2-0162
Atomic Claim 158/574 · 2026-02-16_inferencex-v2-nvidia-blackwell-vs::IX2-0163
SemiAnalysis is free open source software and reader-supported. To receive new posts and support our work, consider becoming a free or paid subscriber.
Atomic Claim 159/574 · 2026-02-16_inferencex-v2-nvidia-blackwell-vs::IX2-0164
Claim: SemiAnalysis 的相關軟體為免費開源,並由讀者支持。
Frame:CLAIM_ONLY· Mode:ASSERTED· Mapping:CLAIM_ONLY
開啟逐條審核
Subscribed

Source: SemiAnalysis InferenceX ↗
Despite the MI355X being competitive in FP8 disagg, its FP4 performance suffers from composability issues. AMD single node FP4 performance is decent, but when we compare AMD FP4 disagg prefill to Nvidia, performance is subpar and the MI355X gets absolutely mogged by Nvidia’s B200. In a 1k1k scenario, the MI355X (MoRI SGLang) with MTP barely manages to beat the B200 (Dynamo SGLang) without MTP.
Atomic Claim 160/574 · 2026-02-16_inferencex-v2-nvidia-blackwell-vs::IX2-0165
Atomic Claim 161/574 · 2026-02-16_inferencex-v2-nvidia-blackwell-vs::IX2-0166
Atomic Claim 162/574 · 2026-02-16_inferencex-v2-nvidia-blackwell-vs::IX2-0167
Atomic Claim 163/574 · 2026-02-16_inferencex-v2-nvidia-blackwell-vs::IX2-0168

Source: SemiAnalysis InferenceX ↗
Once we bring Dynamo TRT-LLM into the equation, the B200’s performance is boosted even more to the point that the MI355X even with MTP can’t match the B200’s performance with Dynamo TRT-LLM and MTP. The MI355X can only match the B200 (without MTP) in performance by using MTP, and only for a range of interactivities from ~60 tok/s/user through ~120 tok/s/user.
Atomic Claim 164/574 · 2026-02-16_inferencex-v2-nvidia-blackwell-vs::IX2-0169
Atomic Claim 165/574 · 2026-02-16_inferencex-v2-nvidia-blackwell-vs::IX2-0170

Source: SemiAnalysis InferenceX ↗
When comparing Dynamo TRTLLM B200 disagg prefill to SGLang MoRI MI355 disagg prefill, AMD gets framemogged due to the more mature implementation of disagg prefill on TRTLLM.
Atomic Claim 166/574 · 2026-02-16_inferencex-v2-nvidia-blackwell-vs::IX2-0171

Source: SemiAnalysis InferenceX ↗

Source: Dwarkesh Podcast and SemiAnalysis
The diagram below shows us the various parallelism configurations that form up the MI355X (MoRI SGLang) Pareto frontier. Note that currently, wide EP is not employed for any points (i.e., configurations with EP 16, 32, etc.).
Atomic Claim 167/574 · 2026-02-16_inferencex-v2-nvidia-blackwell-vs::IX2-0172
Atomic Claim 168/574 · 2026-02-16_inferencex-v2-nvidia-blackwell-vs::IX2-0173

Source: SemiAnalysis InferenceX ↗
Unpacking Inference Providers’ Unit Economics
Below is a list on OpenRouter of all inference providers that serve DeepSeek R1 0528 FP8 along with their cost per million input/output tokens and average interactivity listed on. Disregarding Chutes, the middle of the pack provider serves at an interactivity of around 35 tok/s/user.
Atomic Claim 169/574 · 2026-02-16_inferencex-v2-nvidia-blackwell-vs::IX2-0174
Claim: 文章列出 OpenRouter 上所有提供 DeepSeek R1 0528 FP8 的 inference providers,以及各自每百萬 input/output tokens 價格與平均 interactivity。
Frame:NARY_RELATION· Mode:ASSERTED· Mapping:PARTIAL
開啟逐條審核

Source: OpenRouter ↗
We can then use real InferenceX data to interpolate the cost per million input/output tokens at an interactivity level of 35 tok/sec/user, which is a reasonable interactivity level given the data above.
Atomic Claim 170/574 · 2026-02-16_inferencex-v2-nvidia-blackwell-vs::IX2-0176
Claim: 因此可以利用真實 InferenceX 資料,插值估算在 35 tok/sec/user interactivity 下每百萬 input/output tokens 的成本;根據上述資料,這是一個合理 interactivity 水準。
Frame:ATTRIBUTE· Mode:ASSERTED· Mapping:PARTIAL
開啟逐條審核
As we mention later in the article, this is best understood as _baseline _data and not completely representative of real-world inference, mainly because InferenceX benchmarks on random data and disables prefix caching. In other words, performance/cost will be _at least _this good. It is also important to note that there are not data points for each GPU at _each _interactivity level. Thus we cannot make _exact _comparisons at each degree of interactivity. We nevertheless think the bar chart comparisons presented below are (very) reasonable interpolations in lieu of using exact data points.
Atomic Claim 171/574 · 2026-02-16_inferencex-v2-nvidia-blackwell-vs::IX2-0177
Claim: 如文章後段說明,這些結果更適合視為 baseline,而不完全代表 real-world inference,主要因為 InferenceX 使用隨機資料做 benchmark,且關閉 prefix caching。
Frame:RELATION· Mode:ASSERTED· Mapping:COMPLETE
開啟逐條審核
Atomic Claim 172/574 · 2026-02-16_inferencex-v2-nvidia-blackwell-vs::IX2-0178
Claim: 關於 InferenceX:換句話說,實際 production 的 performance/cost 至少應該不會比這個 baseline 更差。
Frame:ATTRIBUTE· Mode:ASSERTED· Mapping:COMPLETE
開啟逐條審核
Atomic Claim 173/574 · 2026-02-16_inferencex-v2-nvidia-blackwell-vs::IX2-0179
Atomic Claim 174/574 · 2026-02-16_inferencex-v2-nvidia-blackwell-vs::IX2-0180
Claim: 關於 InferenceX:因此無法在每個 interactivity 水準都做完全精確的一對一比較。
Frame:ATTRIBUTE· Mode:ASSERTED· Mapping:COMPLETE
開啟逐條審核
Atomic Claim 175/574 · 2026-02-16_inferencex-v2-nvidia-blackwell-vs::IX2-0181
Claim: 儘管如此,SemiAnalysis 認為下方 bar-chart comparisons 在缺少完整 data points 的情況下,仍屬非常合理的插值。
Frame:CLAIM_ONLY· Mode:ASSERTED· Mapping:CLAIM_ONLY
開啟逐條審核
Comparing disagg+wideEP configs at this interactivity level, we see just how effective distributed inference techniques are when it comes to both perf/TCO and overall throughput. We also see how large scale up domains (like GB300 and GB200 NVL72) absolutely dominate in total throughput per GPU.
Atomic Claim 176/574 · 2026-02-16_inferencex-v2-nvidia-blackwell-vs::IX2-0183
Claim: 也可以看到像 GB300 與 GB200 NVL72 這類大型 scale-up domains,在每 GPU 總 throughput 上具有明顯優勢。
Frame:ATTRIBUTE· Mode:ASSERTED· Mapping:COMPLETE
開啟逐條審核
It is interesting to note that at this interactivity level (on an 8k1k workload type), the B200 can achieve the best perf/TCO when MTP is enabled. Below we also list the Total Cost of Ownership (TCO) (Owning – Hyperscaler) for each GPU:
Atomic Claim 177/574 · 2026-02-16_inferencex-v2-nvidia-blackwell-vs::IX2-0184
Atomic Claim 178/574 · 2026-02-16_inferencex-v2-nvidia-blackwell-vs::IX2-0185
Claim: 文章也列出各 GPU 的 Total Cost of Ownership(TCO,Owning – Hyperscaler)。
Frame:ATTRIBUTE· Mode:ASSERTED· Mapping:COMPLETE
開啟逐條審核

Source: SemiAnalysis TCO Model ↗



Source: SemiAnalysis InferenceX
Let’s use the findings above to dig deeper into the unit economics of serving LLMs at scale. From the OpenRouter data above, we see that Crusoe serves at 36 tok/sec/user at 5.40/M output tokens. If we assume no cache hits and that Crusoe is using at least H200s with SOTA inference techniques like MTP, disagg, and wide EP, the data above suggests they incur a cost of _no more than _/M input tokens and $2.955/M output tokens for a profit margin of up to 83% gross margin (depreciation counted in cost of goods sold) on input tokens and 45% gross margin on output tokens.
Atomic Claim 179/574 · 2026-02-16_inferencex-v2-nvidia-blackwell-vs::IX2-0189
Atomic Claim 180/574 · 2026-02-16_inferencex-v2-nvidia-blackwell-vs::IX2-0186
Atomic Claim 181/574 · 2026-02-16_inferencex-v2-nvidia-blackwell-vs::IX2-0187
Claim: 根據 OpenRouter 資料,Crusoe 以 36 tok/sec/user 提供服務,input token 價格為每百萬 1.35 美元,output token 為每百萬 5.40 美元。
Frame:RELATION· Mode:ASSERTED· Mapping:COMPLETE
開啟逐條審核
Atomic Claim 182/574 · 2026-02-16_inferencex-v2-nvidia-blackwell-vs::IX2-0188
SemiAnalysis InferenceX is free open source software and reader-supported. To receive new posts and support our work, consider becoming a free or paid subscriber.
Subscribed
Of course, these assumptions may not be _exactly _correct and these calculations don’t account for downtime or underutilization, but this gives an idea of some cool math you can do with InferenceX data. More analysis on the economics of inference can be found in the SemiAnalysis Tokenomics Model ↗.
Atomic Claim 183/574 · 2026-02-16_inferencex-v2-nvidia-blackwell-vs::IX2-0190
Claim: 關於 InferenceX:當然,這些假設可能不完全正確,而且計算沒有納入 downtime 或 underutilization。
Frame:ATTRIBUTE· Mode:HYPOTHETICAL· Mapping:COMPLETE
開啟逐條審核
Atomic Claim 184/574 · 2026-02-16_inferencex-v2-nvidia-blackwell-vs::IX2-0191
Claim: 但這可以示範如何利用 InferenceX 資料進行有價值的 unit-economics 計算。
Frame:ATTRIBUTE· Mode:ASSERTED· Mapping:COMPLETE
開啟逐條審核
Atomic Claim 185/574 · 2026-02-16_inferencex-v2-nvidia-blackwell-vs::IX2-0192
Claim: 更多推論經濟性分析可參考 SemiAnalysis Tokenomics Model。
Frame:ATTRIBUTE· Mode:ASSERTED· Mapping:PARTIAL
開啟逐條審核
The OpenRouter data also shows Nebius AI Studio (Fast) serving DeepSeek FP4 at 167 tok/sec/user at 6/M output tokens. Adjusting the interactivity level in InferenceX accordingly and we see the following data.
Atomic Claim 186/574 · 2026-02-16_inferencex-v2-nvidia-blackwell-vs::IX2-0193
Claim: OpenRouter 資料也顯示,Nebius AI Studio(Fast)以 167 tok/sec/user 提供 DeepSeek FP4,input token 每百萬 2 美元、output token 每百萬 6 美元。
Frame:NARY_RELATION· Mode:ASSERTED· Mapping:COMPLETE
開啟逐條審核
Atomic Claim 187/574 · 2026-02-16_inferencex-v2-nvidia-blackwell-vs::IX2-0194
Claim: 將 InferenceX 的 interactivity 調整到對應水準後,可以得到後續資料。
Frame:ATTRIBUTE· Mode:ASSERTED· Mapping:COMPLETE
開啟逐條審核



Source: SemiAnalysis InferenceX
At this high of interactivity, it becomes necessary to employ speculative decoding techniques like MTP to achieve high enough throughput to make inference economical. Luckily, MTP can increase throughput with relatively low risk to overall model accuracy. We will go on to talk more about MTP, and how it can be applied to increase throughput / decrease cost, in later sections of this article.
Atomic Claim 188/574 · 2026-02-16_inferencex-v2-nvidia-blackwell-vs::IX2-0195
Claim: 在這麼高的 interactivity 下,必須採用 MTP 等 speculative decoding 技術,才能讓 throughput 高到足以使推論具經濟效益。
Frame:ATTRIBUTE· Mode:ASSERTED· Mapping:COMPLETE
開啟逐條審核
Atomic Claim 189/574 · 2026-02-16_inferencex-v2-nvidia-blackwell-vs::IX2-0196
Atomic Claim 190/574 · 2026-02-16_inferencex-v2-nvidia-blackwell-vs::IX2-0197
Atomic Claim 191/574 · 2026-02-16_inferencex-v2-nvidia-blackwell-vs::IX2-0198
Lastly, we show one more chart of an FP8 DeepSeek workload served at 125 tok/s/user. This is another low latency workload where MTP considerably improves economic viability. As with the previous example, we note that at these higher ranges of interactivity, the cheapest configs all use MTP.
Atomic Claim 192/574 · 2026-02-16_inferencex-v2-nvidia-blackwell-vs::IX2-0199
Atomic Claim 193/574 · 2026-02-16_inferencex-v2-nvidia-blackwell-vs::IX2-0200
Atomic Claim 194/574 · 2026-02-16_inferencex-v2-nvidia-blackwell-vs::IX2-0201

Source: SemiAnalysis InferenceX ↗
Nvidia Disagg Prefill and WideEP
EP requires all-to-all communication, where every GPU needs to send tokens to every other GPU. This is extremely bandwidth hungry. Recall that Nvidia’s servers have two separate networking domains – the scale-up NVLink domain, and the Scale-out Domain, usually using InfiniBand or Ethernet as the networking protocol.
Atomic Claim 195/574 · 2026-02-16_inferencex-v2-nvidia-blackwell-vs::IX2-0202
Atomic Claim 196/574 · 2026-02-16_inferencex-v2-nvidia-blackwell-vs::IX2-0204
Atomic Claim 197/574 · 2026-02-16_inferencex-v2-nvidia-blackwell-vs::IX2-0205
Claim: 另一個是 Scale-out Domain,通常採用 InfiniBand 或 Ethernet 作為 networking protocol。
Frame:RELATION· Mode:ASSERTED· Mapping:COMPLETE
開啟逐條審核
NVLink domain (within the NVL72 rack): 72 GPUs connected via NVLink with 900 GB/s uni-directional bandwidth per GPU. This is roughly 7-10x the bandwidth of the InfiniBand/Ethernet based scale-out network.
Atomic Claim 198/574 · 2026-02-16_inferencex-v2-nvidia-blackwell-vs::IX2-0206
Atomic Claim 199/574 · 2026-02-16_inferencex-v2-nvidia-blackwell-vs::IX2-0207
Claim: 這大約是基於 InfiniBand/Ethernet 的 scale-out network 頻寬的 7~10 倍。
Frame:COMPARISON· Mode:ASSERTED· Mapping:COMPLETE
開啟逐條審核
InfiniBand/RoCEv2 Ethernet (outside of the NVL72 rack): Typically 400-800 Gbit/s per GPU uni-directional (50-100 GB/s). Note that all our testing for Nvidia was conducted on InfiniBand based clusters.
Atomic Claim 200/574 · 2026-02-16_inferencex-v2-nvidia-blackwell-vs::IX2-0208
Atomic Claim 201/574 · 2026-02-16_inferencex-v2-nvidia-blackwell-vs::IX2-0209
Claim: InferenceX 對 Nvidia 的所有測試都在以 InfiniBand 為基礎的 clusters 上進行。
Frame:ATTRIBUTE· Mode:ASSERTED· Mapping:PARTIAL
開啟逐條審核
TP shards every layer’s weight matrices across GPUs. This means that every single token at every single layer requires up to two all-reduce communications (one after the column-parallel GEMM, one after the row-parallel GEMM). For EP, all-to-all is done only at MoE layers. Each GPU sends only the tokens routed to each expert. This means cheaper comms across all layers for EP vs TP.
Atomic Claim 202/574 · 2026-02-16_inferencex-v2-nvidia-blackwell-vs::IX2-0210
Atomic Claim 203/574 · 2026-02-16_inferencex-v2-nvidia-blackwell-vs::IX2-0211
Claim: 這代表每一層的每個 token 最多需要進行兩次 all-reduce 通訊,一次在 column-parallel GEMM 後、一次在 row-parallel GEMM 後。
Frame:ATTRIBUTE· Mode:ASSERTED· Mapping:COMPLETE
開啟逐條審核
Atomic Claim 204/574 · 2026-02-16_inferencex-v2-nvidia-blackwell-vs::IX2-0212
Atomic Claim 205/574 · 2026-02-16_inferencex-v2-nvidia-blackwell-vs::IX2-0213
Atomic Claim 206/574 · 2026-02-16_inferencex-v2-nvidia-blackwell-vs::IX2-0214
Because EP’s all-to-all communication bandwidth requirements scale with the number of participants, staying within the high-bandwidth NVLink domain before having to cross the slower IB/Eth fabric is better. With NVL72, EP across 72 GPUs is possible without ever leaving NVLink, whereas previous generations (with only 8-GPU NVLink domains) could only do EP across 8 GPUs at NVLink speed before hitting the slower IB/Eth networks.
Atomic Claim 207/574 · 2026-02-16_inferencex-v2-nvidia-blackwell-vs::IX2-0215
Atomic Claim 208/574 · 2026-02-16_inferencex-v2-nvidia-blackwell-vs::IX2-0216
Atomic Claim 209/574 · 2026-02-16_inferencex-v2-nvidia-blackwell-vs::IX2-0217
SemiAnalysis InferenceX is free open source software and reader-supported. To receive new posts and support our work, consider becoming a free or paid subscriber.
Subscribed
Wide EP also has a major advantage in weight loading efficiency. For a model like DeepSeek R1, decode is memory-bandwidth-bound: the bottleneck is how fast GPUs can load weights from HBM. With wide EP (e.g., DEP32), 32 GPUs collectively hold and load the 670B weights once, each loading only its shard (~21B). The total HBM bandwidth of all 32 chips is applied to loading a single copy of the model. By contrast, with narrower EP and more DP replicas (e.g., 5xDEP8), each of the 5 replicas needs its own full copy of the 670B weights, that’s 5×670B = 3.35T of redundant weight loading across the system. EP amortizes weights across chips; DP replicates them. This is why wider EP, enabled by high-bandwidth interconnects like NVLink, delivers significantly better throughput per GPU.
Atomic Claim 210/574 · 2026-02-16_inferencex-v2-nvidia-blackwell-vs::IX2-0218
Atomic Claim 211/574 · 2026-02-16_inferencex-v2-nvidia-blackwell-vs::IX2-0219
Atomic Claim 212/574 · 2026-02-16_inferencex-v2-nvidia-blackwell-vs::IX2-0220
Atomic Claim 213/574 · 2026-02-16_inferencex-v2-nvidia-blackwell-vs::IX2-0221
Atomic Claim 214/574 · 2026-02-16_inferencex-v2-nvidia-blackwell-vs::IX2-0222
Atomic Claim 215/574 · 2026-02-16_inferencex-v2-nvidia-blackwell-vs::IX2-0223
Atomic Claim 216/574 · 2026-02-16_inferencex-v2-nvidia-blackwell-vs::IX2-0225

Source: SemiAnalysis InferenceX ↗
Generally, TP is preferred at lower concurrencies due to load balancing. At small batch sizes, EP suffers from uneven token-to-expert routing, leaving some GPUs underutilized while others are overloaded. TP avoids this since each GPU holds a slice of every expert and always gets an equal share of work. At lower concurrency, the cost of this load imbalance outweighs TP’s additional communication overhead.
Atomic Claim 217/574 · 2026-02-16_inferencex-v2-nvidia-blackwell-vs::IX2-0226
Atomic Claim 218/574 · 2026-02-16_inferencex-v2-nvidia-blackwell-vs::IX2-0227
Atomic Claim 219/574 · 2026-02-16_inferencex-v2-nvidia-blackwell-vs::IX2-0228
Atomic Claim 220/574 · 2026-02-16_inferencex-v2-nvidia-blackwell-vs::IX2-0229
At higher concurrencies, this tradeoff changes. Expert activation becomes more evenly distributed across larger batch sizes, and EP’s communication and weight-loading advantages dominate over TP’s expensive per-layer all-reduce. In the middle of the curve, hybrid TP+EP configurations balance both concerns using small TP groups within each expert for load balancing while EP is used across the wider set of GPUs to amortize weights and reduce communication.
Atomic Claim 221/574 · 2026-02-16_inferencex-v2-nvidia-blackwell-vs::IX2-0231
Claim: 隨 batch size 變大,expert activation 會分布得更平均,此時 EP 在通訊與 weight loading 上的優勢會超越 TP 每層昂貴的 all-reduce 成本。
Frame:ATTRIBUTE· Mode:ASSERTED· Mapping:COMPLETE
開啟逐條審核
Atomic Claim 222/574 · 2026-02-16_inferencex-v2-nvidia-blackwell-vs::IX2-0232
For higher interactivity levels (low batch size), large scale-up world sizes tend not to deliver stronger performance. B300 disagg over IB has the same performance as GB300 with NVL72, since the workload is latency-bound, not bandwidth-bound. The massive NVLink bandwidth advantage of NVL72 doesn’t matter because not even the much slower IB link is saturated by the tiny batches of tokens in flight.
Atomic Claim 223/574 · 2026-02-16_inferencex-v2-nvidia-blackwell-vs::IX2-0233
Claim: 在較高 interactivity、也就是低 batch size 情況下,大型 scale-up world size 通常不會帶來更強效能。
Frame:COMPARISON· Mode:ASSERTED· Mapping:PARTIAL
開啟逐條審核
Atomic Claim 224/574 · 2026-02-16_inferencex-v2-nvidia-blackwell-vs::IX2-0234
Atomic Claim 225/574 · 2026-02-16_inferencex-v2-nvidia-blackwell-vs::IX2-0235
Prefill/decode disaggregation also plays a role. Prefill is compute-heavy and bursty; decode is memory-bandwidth-bound and steady-state. When they share the same GPUs, they interfere with each other, causing latency jitter and wasted capacity. Separating them onto dedicated GPU pools lets each run a workload matched to its characteristics, improving effective utilization. This is why disaggregated B200 configs outperform single-node B200 in the middle of the throughput-interactivity curve. PD separation combined with wider EP across more GPUs over IB amortizes weights more efficiently than cramming both phases onto a single 8-GPU node.
Atomic Claim 226/574 · 2026-02-16_inferencex-v2-nvidia-blackwell-vs::IX2-0236
Atomic Claim 227/574 · 2026-02-16_inferencex-v2-nvidia-blackwell-vs::IX2-0237
Atomic Claim 228/574 · 2026-02-16_inferencex-v2-nvidia-blackwell-vs::IX2-0238
Atomic Claim 229/574 · 2026-02-16_inferencex-v2-nvidia-blackwell-vs::IX2-0239
Atomic Claim 230/574 · 2026-02-16_inferencex-v2-nvidia-blackwell-vs::IX2-0240
Atomic Claim 231/574 · 2026-02-16_inferencex-v2-nvidia-blackwell-vs::IX2-0241
Atomic Claim 232/574 · 2026-02-16_inferencex-v2-nvidia-blackwell-vs::IX2-0242
Side Note: the 10x inference engineers at TogetherAI noticed an pattern for multi-turn traffic where the requirements of first turn prefill is much different from the following turns prefill’s and disaggregrated it leading to better TTFT performance. ↗
Atomic Claim 233/574 · 2026-02-16_inferencex-v2-nvidia-blackwell-vs::IX2-0243

Source: SemiAnalysis InferenceX ↗
Jensen Under Promising and Overdelivering - Hopper vs Blackwell vs Rack Scale NVL72
At GTC 2024, Jensen was on stage promising up to 30x performance gains from H100 to GB200 NVL72, everyone thought it was classic marketing lookmaxxing and would not be achievable in real world. ↗ Many looked to come up with labels for this perceived use of a reality distortion field so they could crack more Jensen Math jokes. Indeed – we did point to the comparison of 30x performance difference between the worst case ↗ for H200 on FP8 to a reasonable case of the GB200 on FP4.
Atomic Claim 234/574 · 2026-02-16_inferencex-v2-nvidia-blackwell-vs::IX2-0244
Claim: GTC 2024 時,Jensen 曾在台上宣稱從 H100 升級到 GB200 NVL72 最多可獲得 30 倍效能提升,當時許多人認為這只是行銷宣傳、現實中難以達成。
Frame:COMPARISON· Mode:ATTRIBUTED· Mapping:COMPLETE
開啟逐條審核
Atomic Claim 235/574 · 2026-02-16_inferencex-v2-nvidia-blackwell-vs::IX2-0246
Atomic Claim 236/574 · 2026-02-16_inferencex-v2-nvidia-blackwell-vs::IX2-0245
Claim: 很多人甚至試圖為這種被認為是「reality distortion field」的宣傳方式取名,以便繼續開 Jensen Math 的玩笑。
Frame:CLAIM_ONLY· Mode:ASSERTED· Mapping:CLAIM_ONLY
開啟逐條審核

Nvidia Blackwell Perf TCO Analysis - B100 vs B200 vs GB200NVL72
Dylan Patel ↗ and Daniel Nishball ↗
Atomic Claim 237/574 · 2026-02-16_inferencex-v2-nvidia-blackwell-vs::IX2-0247
Claim: 此處署名為 Dylan Patel 與 Daniel Nishball。
Frame:CLAIM_ONLY· Mode:ASSERTED· Mapping:CLAIM_ONLY
開啟逐條審核
·
2024年4月10日
Read full story ↗

But it turns out the joke is on them. Fast forward almost two years later, and we can now see that it wasn’t marketing hype lookmaxing after all, and Jensen was actually under promising on Blackwell performance the whole time. From our testing, Blackwell is so good at large scale MoE inferencing compared to even a strong H100 disagg+wideEP FP8 baseline that it, at 116 toks/s/user, delivers up to 98x better perf on GB200 NVL72 FP4 and up to 100x better perf on GB300 NVL72 FP4! Maybe the new Jensen Math rule is that he delivers double whatever he promises in terms of token throughput. The more you spend, the more you save indeed!
Atomic Claim 238/574 · 2026-02-16_inferencex-v2-nvidia-blackwell-vs::IX2-0248
Claim: 但後來看來,真正被打臉的是當時嘲笑這些數字的人。
Frame:CLAIM_ONLY· Mode:ASSERTED· Mapping:CLAIM_ONLY
開啟逐條審核
Atomic Claim 239/574 · 2026-02-16_inferencex-v2-nvidia-blackwell-vs::IX2-0251
Atomic Claim 240/574 · 2026-02-16_inferencex-v2-nvidia-blackwell-vs::IX2-0252
Claim: 根據 InferenceX 測試,Blackwell 在大規模 MoE 推論上相較即使是很強的 H100 disagg+wideEP FP8 baseline 仍有巨大優勢;在 116 toks/s/user 下,GB200 NVL72 FP4 最高可達 98 倍,GB300 NVL72 FP4 最高可達 100 倍效能。
Frame:COMPARISON· Mode:ASSERTED· Mapping:PARTIAL
開啟逐條審核
Atomic Claim 241/574 · 2026-02-16_inferencex-v2-nvidia-blackwell-vs::IX2-0253
Claim: SemiAnalysis 戲稱新的 Jensen Math 規則可能是:token throughput 最後會做到承諾值的兩倍。
Frame:CLAIM_ONLY· Mode:ASSERTED· Mapping:CLAIM_ONLY
開啟逐條審核

Source: SemiAnalysis InferenceX ↗
Even when factoring in the increased total cost of ownership of Blackwell and Blackwell Ultra, we see a 9.7x(40 tok/s/user) up to 65x(116 tok/s/user) improvement in tokens per dollar compared to Hopper. You can explore Hopper vs Blackwell performance in detail on our free website ↗. Blackwell performance is so good compared to Hopper that we needed to an log scale to our dashboard in order to visualize it.
Atomic Claim 242/574 · 2026-02-16_inferencex-v2-nvidia-blackwell-vs::IX2-0255
Atomic Claim 243/574 · 2026-02-16_inferencex-v2-nvidia-blackwell-vs::IX2-0256
Atomic Claim 244/574 · 2026-02-16_inferencex-v2-nvidia-blackwell-vs::IX2-0257

Source: SemiAnalysis InferenceX ↗
As mentioned earlier in the article, B300 servers only connect at most 8 GPUs using the 900GByte/s/GPU NVLink scale-up network whereas GB300 NVL72 servers connect 72 GPUs using the NVlink scale-up network. So when we need more than 8 GPUs (but less than 72 GPUs) for the inference setup, we need to bring in multiple nodes of B300 servers to form our inference system which means communications falls back to the lower InfiniBand XDR scale-out network featuring 800Gbit/s (uni-di) per GPU of bandwidth. Compare this to a rack scale GB300 NVL72 which connects 72 GPUs over NVLink delivering 900GByte/s (uni-di) per GPU of bandwidth and we can see that the rack-scale server allows the GPUs in the inference setup to talk to each other with over 9x higher bandwidth compared to the case of the multiple nodes of B300 servers.
Atomic Claim 245/574 · 2026-02-16_inferencex-v2-nvidia-blackwell-vs::IX2-0258
Claim: 如前文所述,B300 server 最多只能用 900GByte/s/GPU 的 NVLink scale-up network 連接 8 顆 GPUs,而 GB300 NVL72 server 可在同一 NVlink scale-up network 內連接 72 顆 GPUs。
Frame:NARY_RELATION· Mode:ASSERTED· Mapping:PARTIAL
開啟逐條審核
Atomic Claim 246/574 · 2026-02-16_inferencex-v2-nvidia-blackwell-vs::IX2-0259
Claim: 因此當推論配置需要超過 8 顆 GPUs、但少於 72 顆 GPUs 時,B300 必須使用多個 nodes 組成系統,通訊就會退回頻寬較低的 InfiniBand XDR scale-out network,每顆 GPU 僅有 800Gbit/s uni-di。
Frame:COMPARISON· Mode:ASSERTED· Mapping:PARTIAL
開啟逐條審核
Atomic Claim 247/574 · 2026-02-16_inferencex-v2-nvidia-blackwell-vs::IX2-0260
SemiAnalysis is free open source software and reader-supported. To receive new posts and support our work, consider becoming a free or paid subscriber.
Atomic Claim 248/574 · 2026-02-16_inferencex-v2-nvidia-blackwell-vs::IX2-0261
Claim: SemiAnalysis 的相關軟體為免費開源,並由讀者支持。
Frame:CLAIM_ONLY· Mode:ASSERTED· Mapping:CLAIM_ONLY
開啟逐條審核
Subscribed

Source: SemiAnalysis InferenceX ↗
Admittedly the GB300 NVL72 has a higher all-in cost per GPU, but this only reduces the bandwidth per TCO advantage to being 8x faster. The bandwidth advantage of the rack-scale architecture directly drives a much lower cost per token. Google TPU, AWS Trainium and Nvidia are the only AI chips to have rack scale system designs deployed today. Engineering samples and low volume production of AMD’s first rack scale MI455X UALoE72 system will be in H2 2026 while due to manufacturing delays, the mass production ramp and first production tokens will only be generated on an MI455X UALoE72 by Q2 2027.
Atomic Claim 249/574 · 2026-02-16_inferencex-v2-nvidia-blackwell-vs::IX2-0262
Claim: 雖然 GB300 NVL72 每顆 GPU 的 all-in cost 較高,但折算成 bandwidth per TCO 後,仍有約 8 倍速度優勢。
Frame:COMPARISON· Mode:ASSERTED· Mapping:COMPLETE
開啟逐條審核
Atomic Claim 250/574 · 2026-02-16_inferencex-v2-nvidia-blackwell-vs::IX2-0263
Claim: rack-scale architecture 的頻寬優勢會直接轉化成更低的每 token 成本。
Frame:ATTRIBUTE· Mode:ASSERTED· Mapping:PARTIAL
開啟逐條審核
Atomic Claim 251/574 · 2026-02-16_inferencex-v2-nvidia-blackwell-vs::IX2-0264
Atomic Claim 252/574 · 2026-02-16_inferencex-v2-nvidia-blackwell-vs::IX2-0265
Claim: AMD MI455X UALoE72 的 engineering samples 與低量生產預計在 2026 年下半年開始。
Frame:ATTRIBUTE· Mode:ASSERTED· Mapping:PARTIAL
開啟逐條審核
Atomic Claim 253/574 · 2026-02-16_inferencex-v2-nvidia-blackwell-vs::IX2-0266
Claim: 由於製造延遲,MI455X UALoE72 mass production 預計要到 2027 年第二季。
Frame:ATTRIBUTE· Mode:ASSERTED· Mapping:COMPLETE
開啟逐條審核
Atomic Claim 254/574 · 2026-02-16_inferencex-v2-nvidia-blackwell-vs::IX2-0267
Claim: MI455X UALoE72 真正產出 production tokens 預計也要到 2027 年第二季。
Frame:ATTRIBUTE· Mode:ASSERTED· Mapping:COMPLETE
開啟逐條審核

Source: SemiAnalysis InferenceX ↗
Blackwell vs Blackwell Ultra
On paper, the newly released Blackwell Ultra has the same memory bandwidth as Blackwell, the same FP8 performance and only 1.5x higher FP4 performance, but when measuring we actually see up to 1.5x better FP8 performance on the Blackwell Ultra, though we only see 1.1x better performance on FP4. This may be due to Blackwell Ultra being a newly released GPU, meaning software is not fully optimized yet.
Atomic Claim 255/574 · 2026-02-16_inferencex-v2-nvidia-blackwell-vs::IX2-0270
Atomic Claim 256/574 · 2026-02-16_inferencex-v2-nvidia-blackwell-vs::IX2-0272
Atomic Claim 257/574 · 2026-02-16_inferencex-v2-nvidia-blackwell-vs::IX2-0271
Atomic Claim 258/574 · 2026-02-16_inferencex-v2-nvidia-blackwell-vs::IX2-0269
Atomic Claim 259/574 · 2026-02-16_inferencex-v2-nvidia-blackwell-vs::IX2-0268
Atomic Claim 260/574 · 2026-02-16_inferencex-v2-nvidia-blackwell-vs::IX2-0273


Source: SemiAnalysis InferenceX
MI355X vs MI325X vs MI300X
On AMD SKUs, we see up to 10x better performance on the MI355X vs the MI300X. AMD has only gotten DeepSeek SGLang Disaggregated Inferencing to work on the MI355X so far AMD has not submitted MI300X or MI325X disaggregated inferencing results, potentially due to software issues on older SKUs that are still being solved.
Atomic Claim 261/574 · 2026-02-16_inferencex-v2-nvidia-blackwell-vs::IX2-0274
Atomic Claim 262/574 · 2026-02-16_inferencex-v2-nvidia-blackwell-vs::IX2-0275

Source: SemiAnalysis InferenceX ↗


Source: SemiAnalysis InferenceX
Turning to cost, for DeepSeekR1 on FP8, at an interactivity of 24 tok/s/user, the MI355X delivers inferences a cost that is slightly less than 3x cheaper than for the MI325X. The throughput of each GPU is slightly less than 4 times that of MI325X.
Atomic Claim 263/574 · 2026-02-16_inferencex-v2-nvidia-blackwell-vs::IX2-0276
Claim: 就成本而言,在 DeepSeekR1 FP8、24 tok/s/user interactivity 下,MI355X 的推論成本略低於 MI325X 的三分之一。
Frame:COMPARISON· Mode:ASSERTED· Mapping:PARTIAL
開啟逐條審核
Atomic Claim 264/574 · 2026-02-16_inferencex-v2-nvidia-blackwell-vs::IX2-0277

Source: SemiAnalysis InferenceX ↗
AMD Composability Issue on FP4, Distributed Inferencing and Wide Expert Parallelism
While AMD performs somewhat decently on single node FP4 and performs competitively to B200 SGLang on FP8 distributed inferencing, the issue with the current AMD open source inferencing stack is that, while individual inference optimizations perform well, real customers deploy with multiple optimizations composed together. Top tier AI labs are all using FP4 **with **disaggregated inferencing with wide expert parallelism all enabled at the same time, and this is where the issue occurs.
Atomic Claim 265/574 · 2026-02-16_inferencex-v2-nvidia-blackwell-vs::IX2-0278
Atomic Claim 266/574 · 2026-02-16_inferencex-v2-nvidia-blackwell-vs::IX2-0280
Claim: 頂級 AI labs 已經同時啟用 FP4、disaggregated inferencing 與 wide expert parallelism。
Frame:ATTRIBUTE· Mode:ASSERTED· Mapping:PARTIAL
開啟逐條審核
Atomic Claim 267/574 · 2026-02-16_inferencex-v2-nvidia-blackwell-vs::IX2-0281
SemiAnalysis is free open source software and reader-supported. To receive new posts and support our work, consider becoming a free or paid subscriber.
Atomic Claim 268/574 · 2026-02-16_inferencex-v2-nvidia-blackwell-vs::IX2-0282
Claim: SemiAnalysis 的相關軟體為免費開源,並由讀者支持。
Frame:CLAIM_ONLY· Mode:ASSERTED· Mapping:CLAIM_ONLY
開啟逐條審核
Subscribed
AMD software is still not meeting the mark, and the theoretical speed of light modelling at SemiAnalysis and at AMD show that for FP4, disaggregated inferencing with wide expert parallelism should perform better than inference on a single node of MI355X. Unfortunately, Software continues to be a massive bottleneck for AMD GPUs. AMD management needs to continue to sharpen resource allocation of their engineering talent, for instance, re-allocate their engineering resources away from pet single node projects that nobody uses like ATOM towards fixing the aforementioned issues with composability of inference optimizations between disaggregated inferencing, wide expert parallelism and FP4. The current subpar software is due to lack of focus and incorrect prioritization of where the industry already is at. All top tier labs are already using disaggregated inferencing and wide expert parallelism; AMD needs to stop focusing on single node and heavily invest focus into multi node inferencing for open source solutions.
Atomic Claim 269/574 · 2026-02-16_inferencex-v2-nvidia-blackwell-vs::IX2-0283
Atomic Claim 270/574 · 2026-02-16_inferencex-v2-nvidia-blackwell-vs::IX2-0284
Claim: SemiAnalysis 與 AMD 的 theoretical speed-of-light modeling 都顯示,在 FP4 下,搭配 wide expert parallelism 的 disaggregated inference 理論上應該優於單一 MI355X node。
Frame:COMPARISON· Mode:EXPECTED· Mapping:COMPLETE
開啟逐條審核
Atomic Claim 271/574 · 2026-02-16_inferencex-v2-nvidia-blackwell-vs::IX2-0285
Atomic Claim 272/574 · 2026-02-16_inferencex-v2-nvidia-blackwell-vs::IX2-0286
Claim: SemiAnalysis 認為 AMD management 應進一步改善工程人才配置,例如把資源從幾乎沒有客戶使用的 single-node pet projects(如 ATOM)轉向解決 disaggregated inference、wide expert parallelism 與 FP4 之間的 composability 問題。
Frame:NARY_RELATION· Mode:ASSERTED· Mapping:COMPLETE
開啟逐條審核
Atomic Claim 273/574 · 2026-02-16_inferencex-v2-nvidia-blackwell-vs::IX2-0288
Claim: 所有頂級 labs 都已使用 disaggregated inferencing 與 wide expert parallelism。
Frame:ATTRIBUTE· Mode:ASSERTED· Mapping:PARTIAL
開啟逐條審核
Atomic Claim 274/574 · 2026-02-16_inferencex-v2-nvidia-blackwell-vs::IX2-0289
AMD is more than six months behind on open source distributed inferencing and wide expert parallelism and FP4 composability as shown by Nvidia and SGLang team showing off their NVFP4 performance on DeepSeek six months ago ↗.
Atomic Claim 275/574 · 2026-02-16_inferencex-v2-nvidia-blackwell-vs::IX2-0290

Source: SemiAnalysis InferenceX ↗
AMD ATOM Engine
AMD has launched a new inference engine called ATOM. Atom can deliver slightly better single node performance, but it is completely lacking on a lot of features that makes it unusable for real workloads. One such example is that it does not support NVMe or CPU KVCache offloading, tool parsing, wide expert parallelism, or disaggregated serving. This has led to zero customers using it in production. Unlike Nvidia’s TRTLLM which generates billions of tokens per hour globally at companies like TogetherAI, etc and does support tool parsing and other features ↗, there are no token factories currently using ATOM due to the lack of the aforementioned features.
Atomic Claim 276/574 · 2026-02-16_inferencex-v2-nvidia-blackwell-vs::IX2-0296
Atomic Claim 277/574 · 2026-02-16_inferencex-v2-nvidia-blackwell-vs::IX2-0291
Claim: AMD 推出一套名為 ATOM 的新 inference engine。
Frame:NARY_RELATION· Mode:ASSERTED· Mapping:COMPLETE
開啟逐條審核
Atomic Claim 278/574 · 2026-02-16_inferencex-v2-nvidia-blackwell-vs::IX2-0292
Atomic Claim 279/574 · 2026-02-16_inferencex-v2-nvidia-blackwell-vs::IX2-0293
Atomic Claim 280/574 · 2026-02-16_inferencex-v2-nvidia-blackwell-vs::IX2-0294
Claim: 例如 ATOM 不支援 NVMe 或 CPU KVCache offloading、tool parsing、wide expert parallelism 或 disaggregated serving。
Frame:NARY_RELATION· Mode:ASSERTED· Mapping:COMPLETE
開啟逐條審核
Atomic Claim 281/574 · 2026-02-16_inferencex-v2-nvidia-blackwell-vs::IX2-0295
Furthermore, maintainers of open-source inference engines like vLLM are disappointed in AMD due to a lack of engineering and GPU resources provided by AMD. For example, Simon Mo, lead vLLM maintainer, states in this GitHub RFC that there is still no working MI355X that he can add to vLLM CI, hence the poor user experience. There are currently zero Mi355X tests on vLLM, while NVIDIA’s B200 has many tests on vLLM. Similarly, there are still not enough MI300X CI machines on vLLM. Upstream vLLM needs at least 20 more MI300 machines, 20 more MI325 machines and 20 more MI355X machines to reach the same level of usability as CUDA.
Atomic Claim 282/574 · 2026-02-16_inferencex-v2-nvidia-blackwell-vs::IX2-0297
Atomic Claim 283/574 · 2026-02-16_inferencex-v2-nvidia-blackwell-vs::IX2-0298
Atomic Claim 284/574 · 2026-02-16_inferencex-v2-nvidia-blackwell-vs::IX2-0299
Atomic Claim 285/574 · 2026-02-16_inferencex-v2-nvidia-blackwell-vs::IX2-0300
Atomic Claim 286/574 · 2026-02-16_inferencex-v2-nvidia-blackwell-vs::IX2-0301
Atomic Claim 287/574 · 2026-02-16_inferencex-v2-nvidia-blackwell-vs::IX2-0302
We at SemiAnalysis have been trying to get AMD to contribute more compute to vLLM and have had some success on that within the couple weeks. vLLM will start to get a couple of MI355X machines such that they can bring their CI test parity from 0% to non-0%. We will talk more about AMD’s previous lackluster contribution towards vLLM, SGLang, PyTorch CI machine situation & how Anush started to fix it in our upcoming State of AMD article. At SemiAnalysis, we will have internal dashboard to track the # of tests & quality of tests that AMD & NVIDIA runs on vLLM, SGLang, PyTorch, & JAX.
Atomic Claim 288/574 · 2026-02-16_inferencex-v2-nvidia-blackwell-vs::IX2-0303
Atomic Claim 289/574 · 2026-02-16_inferencex-v2-nvidia-blackwell-vs::IX2-0304
Atomic Claim 290/574 · 2026-02-16_inferencex-v2-nvidia-blackwell-vs::IX2-0305
Atomic Claim 291/574 · 2026-02-16_inferencex-v2-nvidia-blackwell-vs::IX2-0306
Moreover, the vLLM maintainers say that they cannot support day 0 vLLM support for ROCm due to this issue of lack of machine resources. This huge disparity in time to market continues to lead to ROCm lagging behind and leaving a huge opening for Nvidia to continue to charge an insane 75% gross margin (4x markup on cost of goods).
Atomic Claim 292/574 · 2026-02-16_inferencex-v2-nvidia-blackwell-vs::IX2-0307
Atomic Claim 293/574 · 2026-02-16_inferencex-v2-nvidia-blackwell-vs::IX2-0308

Lastly, AMD has not had enough committers “who demonstrated sustained upstream engagement through feature shepherding and code ownership” and has a lack of reviewers that can review their own code. This is why the pace of development on ROCm vLLM has been much slower than for CUDA vLLM.
Atomic Claim 294/574 · 2026-02-16_inferencex-v2-nvidia-blackwell-vs::IX2-0309
Atomic Claim 295/574 · 2026-02-16_inferencex-v2-nvidia-blackwell-vs::IX2-0310
Atomic Claim 296/574 · 2026-02-16_inferencex-v2-nvidia-blackwell-vs::IX2-0311
There are many talented 10x engineers at AMD that work on ATOM and we would encourage AMD management to think about re-deploying these 10x engineers towards working on libraries and frameworks that people actually use, such as vLLM and SGLang.
Atomic Claim 297/574 · 2026-02-16_inferencex-v2-nvidia-blackwell-vs::IX2-0312
As we mentioned earlier, AMD also needs to prioritize addressing composability issues with FP4, wideEP and disaggregated serving as opposed to overly focusing on optimizing FP4 for a single node.
Atomic Claim 298/574 · 2026-02-16_inferencex-v2-nvidia-blackwell-vs::IX2-0313
Claim: 如前文所述,AMD 也應優先解決 FP4、wideEP 與 disaggregated serving 的 composability 問題,而不是過度聚焦 single-node FP4 最佳化。
Frame:ATTRIBUTE· Mode:ASSERTED· Mapping:COMPLETE
開啟逐條審核

Source: SemiAnalysis InferenceX ↗
Multi Token Prediction (MTP)
Speculative decoding reduces the cost of autoregressive generation by using a small, inexpensive draft model to propose several tokens ahead. The large model then checks the proposed tokens in a single forward pass that resembles a prefill computation. For a given input sequence length, a single forward pass can take roughly the same time when the input has N more tokens. Speculative decoding uses this property to run inference on a smaller model to draft multiple tokens for the main model to verify with a single forward pass, producing at most N additional tokens in a similar time budget.
Atomic Claim 299/574 · 2026-02-16_inferencex-v2-nvidia-blackwell-vs::IX2-0314
Claim: Speculative decoding 透過較小、成本較低的 draft model 預先提出多個 token,降低 autoregressive generation 成本。
Frame:COMPARISON· Mode:ASSERTED· Mapping:PARTIAL
開啟逐條審核
Atomic Claim 300/574 · 2026-02-16_inferencex-v2-nvidia-blackwell-vs::IX2-0315
Atomic Claim 301/574 · 2026-02-16_inferencex-v2-nvidia-blackwell-vs::IX2-0316
Claim: 對固定輸入 sequence length 而言,即使輸入多 N 個 token,一次 forward pass 所需時間可能仍大致相同。
Frame:COMPARISON· Mode:ASSERTED· Mapping:PARTIAL
開啟逐條審核
Atomic Claim 302/574 · 2026-02-16_inferencex-v2-nvidia-blackwell-vs::IX2-0317
Claim: Speculative decoding 利用此特性,先以較小模型一次 draft 多個 tokens,再讓主模型用單次 forward pass 驗證,在近似相同時間預算內最多產生 N 個額外 token。
Frame:COMPARISON· Mode:ASSERTED· Mapping:PARTIAL
開啟逐條審核

Source: Brendan Bycroft ↗
This assumption regarding additional token production with the same time budget is strongest for dense models because batched verification can reuse the same weight stream across multiple positions. For Mixture-of-Experts models, different tokens may route to different experts, so verifying multiple draft tokens can activate more experts than single-token decoding and force additional expert weights to be fetched from memory. As shown in the Mixtral 8x7B Instruct model results in the EAGLE paper, this extra memory traffic erodes bandwidth savings and can make verification notably comparable to a standard decoding step.
Atomic Claim 303/574 · 2026-02-16_inferencex-v2-nvidia-blackwell-vs::IX2-0318
Claim: 「在相同時間內產生更多 token」的假設對 dense models 最成立,因為 batched verification 可以在多個 positions 重複利用相同 weight stream。
Frame:COMPARISON· Mode:HYPOTHETICAL· Mapping:PARTIAL
開啟逐條審核
Atomic Claim 304/574 · 2026-02-16_inferencex-v2-nvidia-blackwell-vs::IX2-0319
Claim: 對 Mixture-of-Experts models,不同 token 可能 route 到不同 experts,因此一次驗證多個 draft tokens 可能啟動更多 experts,迫使系統從記憶體載入額外 expert weights。
Frame:ATTRIBUTE· Mode:HYPOTHETICAL· Mapping:COMPLETE
開啟逐條審核
Atomic Claim 305/574 · 2026-02-16_inferencex-v2-nvidia-blackwell-vs::IX2-0320
Claim: EAGLE 論文中的 Mixtral 8x7B Instruct 結果顯示,額外 memory traffic 會侵蝕 bandwidth savings,使 verification 成本可能接近一般 decoding step。
Frame:ATTRIBUTE· Mode:ASSERTED· Mapping:COMPLETE
開啟逐條審核
Multi-token prediction pursues similar benefits without requiring a separate draft model. Auxiliary prediction heads are added to the model architecture, so a single model can propose several future tokens from the same underlying representation. This improves distribution alignment because the proposals come from the same model that ultimately scores them. Multi-token prediction also avoids the operational complexity of serving an additional model while still enabling multi-token generation strategies but requires the MTP heads to be pretrained alongside the main model.
Atomic Claim 306/574 · 2026-02-16_inferencex-v2-nvidia-blackwell-vs::IX2-0321
Claim: Multi-token prediction 追求類似效益,但不需要額外的 draft model。
Frame:ATTRIBUTE· Mode:ASSERTED· Mapping:PARTIAL
開啟逐條審核
Atomic Claim 307/574 · 2026-02-16_inferencex-v2-nvidia-blackwell-vs::IX2-0324
Claim: Multi-token prediction 不需要額外 serving 一個模型,因此可避免操作複雜度,同時支援 multi-token generation;但 MTP heads 必須與主模型一起預訓練。
Frame:ATTRIBUTE· Mode:ASSERTED· Mapping:COMPLETE
開啟逐條審核

Source: SemiAnalysis InferenceX ↗
Across all SKUs, enabling MTP results in performance gains. By making use of the typically unused logits to verify the extra tokens, minimal compute overhead is added, saving extra expensive weight loads during decode.
Atomic Claim 308/574 · 2026-02-16_inferencex-v2-nvidia-blackwell-vs::IX2-0325
Atomic Claim 309/574 · 2026-02-16_inferencex-v2-nvidia-blackwell-vs::IX2-0326

Source: SemiAnalysis InferenceX ↗
At large batch sizes, the inference regime is less memory-bandwidth bound compared to for low batch sizes. Since speculative decoding (including MTP) works by trading excess compute for fewer memory-bound decoding steps, this extra verification work from speculative tokens may not fit cleanly into slack, resulting in smaller improvements at high batch sizes.
Atomic Claim 310/574 · 2026-02-16_inferencex-v2-nvidia-blackwell-vs::IX2-0327
Claim: 在大 batch size 下,推論較不受記憶體頻寬限制。
Frame:ATTRIBUTE· Mode:ASSERTED· Mapping:COMPLETE
開啟逐條審核
Atomic Claim 311/574 · 2026-02-16_inferencex-v2-nvidia-blackwell-vs::IX2-0328
Claim: 由於 speculative decoding(包含 MTP)本質上是用額外 compute 換取更少 memory-bound decoding steps,因此在高 batch size 下,speculative tokens 帶來的額外 verification 工作未必能完全塞進閒置 compute,效益可能較小。
Frame:COMPARISON· Mode:HYPOTHETICAL· Mapping:COMPLETE
開啟逐條審核
In terms of cost, MTP can drive huge cost savings, in the below table, we see that DeepSeek-R1-0528 run on FP4 using Dynamo TRT costs 0.057 per million total tokens.
Atomic Claim 312/574 · 2026-02-16_inferencex-v2-nvidia-blackwell-vs::IX2-0329
Atomic Claim 313/574 · 2026-02-16_inferencex-v2-nvidia-blackwell-vs::IX2-0330

Source: SemiAnalysis InferenceX ↗
In all configs, when all else is held equal, using MTP with DeepSeek R1 increases interactivity with no significant impact on model accuracy. This is in line with the DeepSeek V3 tech report findings.
Atomic Claim 314/574 · 2026-02-16_inferencex-v2-nvidia-blackwell-vs::IX2-0331
Atomic Claim 315/574 · 2026-02-16_inferencex-v2-nvidia-blackwell-vs::IX2-0332

Source: SemiAnalysis InferenceX ↗
Regarding the validity of MTP performance numbers, one may argue that the distribution of a synthetic dataset may not resemble real data. However, comparing MTP acceptance behavior between MTBench and our 1k1k benchmark, we see a very similar distribution confirming that our InferenceX benchmark is a good proxy for real world production performance. That said, InferenceX is not perfect and we are always looking to improve. If you want to be part of the mission, apply to join our special projects team here ↗.
Atomic Claim 316/574 · 2026-02-16_inferencex-v2-nvidia-blackwell-vs::IX2-0333
Atomic Claim 317/574 · 2026-02-16_inferencex-v2-nvidia-blackwell-vs::IX2-0334
Claim: 但比較 MTBench 與 InferenceX 1k1k benchmark 的 MTP acceptance behavior,可以看到非常相似的分布,支持 InferenceX benchmark 可作為 real-world production performance 的良好 proxy。
Frame:COMPARISON· Mode:ASSERTED· Mapping:COMPLETE
開啟逐條審核
Atomic Claim 318/574 · 2026-02-16_inferencex-v2-nvidia-blackwell-vs::IX2-0335
Claim: 不過 InferenceX 並不完美,團隊仍持續尋找改善方式。
Frame:ATTRIBUTE· Mode:ASSERTED· Mapping:PARTIAL
開啟逐條審核
Atomic Claim 319/574 · 2026-02-16_inferencex-v2-nvidia-blackwell-vs::IX2-0336
Claim: 原文此處邀請有興趣者加入其 special projects team。
Frame:CLAIM_ONLY· Mode:ASSERTED· Mapping:CLAIM_ONLY
開啟逐條審核

Source: SemiAnalysis InferenceX ↗
Accuracy Evaluations
Throughput optimizations can sometimes quietly trade off accuracy (e.g. via aggressively relaxed acceptance rates, decoding tweaks, numerically unstable kernels, or endpoint misconfiguration). Without evals, a misconfigured server (truncation, bad decoding, wrong endpoint params) can still produce great throughput numbers but deliver garbage answers. For example, this additional layer of checks has helped us discover issues with some DP attention implementation for GPT-OSS.
Atomic Claim 320/574 · 2026-02-16_inferencex-v2-nvidia-blackwell-vs::IX2-0337
Claim: 關於 InferenceX:Throughput optimizations 有時可能在不明顯的情況下犧牲 accuracy,例如過度放寬 acceptance rates、修改 decoding、使用數值不穩定 kernels,或 endpoint misconfiguration。
Frame:ATTRIBUTE· Mode:ASSERTED· Mapping:COMPLETE
開啟逐條審核
Atomic Claim 321/574 · 2026-02-16_inferencex-v2-nvidia-blackwell-vs::IX2-0338
Claim: 關於 InferenceX:若沒有 evals,即使 server configuration 錯誤,例如 truncation、bad decoding 或 endpoint params 錯誤,也可能得到漂亮 throughput 數字,但輸出品質非常差。
Frame:ATTRIBUTE· Mode:HYPOTHETICAL· Mapping:COMPLETE
開啟逐條審核
Atomic Claim 322/574 · 2026-02-16_inferencex-v2-nvidia-blackwell-vs::IX2-0339
Each representative throughput config now has an associated numerical accuracy check. Currently we are only using GSM8k, but being a very easy benchmark, the evaluation scores may not change much from differences in numerical calculation, and a harder benchmark may have a larger delta with respect to numerical accuracy. Thus, we plan to expand towards harder ones in the future, such as GPQA, HLE, MATH-500, SWE-Bench verified.
Atomic Claim 323/574 · 2026-02-16_inferencex-v2-nvidia-blackwell-vs::IX2-0340
Claim: 關於 InferenceX:現在每個代表性的 throughput configuration 都會搭配一項 numerical accuracy check。
Frame:ATTRIBUTE· Mode:ASSERTED· Mapping:COMPLETE
開啟逐條審核
Atomic Claim 324/574 · 2026-02-16_inferencex-v2-nvidia-blackwell-vs::IX2-0341
Claim: 關於 InferenceX:目前 accuracy check 只使用 GSM8k。
Frame:ATTRIBUTE· Mode:ASSERTED· Mapping:PARTIAL
開啟逐條審核
Atomic Claim 325/574 · 2026-02-16_inferencex-v2-nvidia-blackwell-vs::IX2-0342
Claim: 關於 InferenceX:由於 GSM8k 很容易,數值計算上的差異可能不會明顯反映在 evaluation score 上。
Frame:ATTRIBUTE· Mode:HYPOTHETICAL· Mapping:PARTIAL
開啟逐條審核
Atomic Claim 326/574 · 2026-02-16_inferencex-v2-nvidia-blackwell-vs::IX2-0343
Claim: 關於 InferenceX:更困難的 benchmark 在 numerical accuracy 上可能呈現更大的 delta。
Frame:ATTRIBUTE· Mode:HYPOTHETICAL· Mapping:PARTIAL
開啟逐條審核
Atomic Claim 327/574 · 2026-02-16_inferencex-v2-nvidia-blackwell-vs::IX2-0344
Claim: 因此未來計畫加入更難的 benchmarks,例如 GPQA、HLE、MATH-500 與 SWE-Bench verified。
Frame:ATTRIBUTE· Mode:ASSERTED· Mapping:PARTIAL
開啟逐條審核
Another form of performance-accuracy tradeoff is quantization. Serving models at lower precision may result in worse model outputs. For DeepSeek R1, FP8 runs have very slightly higher evaluation scores than FP4. Note that GSM8k evals are saturated and often during QAT/PAT it is calibrated to common popular GSM8k, MATH-500, etc, leading to sometimes evals showing great results while real world end user evaluation being subpar. If we want to be part of the team to figure out how to properly evaluate inference engine accuracy, apply to join the mission here ↗.
Atomic Claim 328/574 · 2026-02-16_inferencex-v2-nvidia-blackwell-vs::IX2-0345
Claim: 另一種 performance-accuracy trade-off 是 quantization。
Frame:ATTRIBUTE· Mode:ASSERTED· Mapping:PARTIAL
開啟逐條審核
Atomic Claim 329/574 · 2026-02-16_inferencex-v2-nvidia-blackwell-vs::IX2-0346
Claim: 關於 InferenceX:以較低 precision serving models 可能導致模型輸出品質變差。
Frame:COMPARISON· Mode:HYPOTHETICAL· Mapping:PARTIAL
開啟逐條審核
Atomic Claim 330/574 · 2026-02-16_inferencex-v2-nvidia-blackwell-vs::IX2-0347
Atomic Claim 331/574 · 2026-02-16_inferencex-v2-nvidia-blackwell-vs::IX2-0348
Claim: 關於 InferenceX:需要注意,GSM8k eval 已高度飽和,而且 QAT/PAT 常針對 GSM8k、MATH-500 等熱門 benchmark 校準,因此 eval 看起來可能很好,但 real-world end-user evaluation 仍可能較差。
Frame:ATTRIBUTE· Mode:ASSERTED· Mapping:PARTIAL
開啟逐條審核
Atomic Claim 332/574 · 2026-02-16_inferencex-v2-nvidia-blackwell-vs::IX2-0349
Claim: 原文此處邀請有興趣者加入團隊,共同研究如何正確評估 inference engine accuracy。
Frame:ATTRIBUTE· Mode:ASSERTED· Mapping:PARTIAL
開啟逐條審核

Source: SemiAnalysis InferenceX ↗
Anthropic Fast Mode Inferencing Explained
Anthropic recently released “fast mode ↗” alongside Opus 4.6. The value proposition: the same model quality at roughly 2.5× the speed, for around 6–12× the price. Both figures might seem surprising, and some users have speculated that this must require new hardware ↗. It doesn’t. In fact, this is just the fundamental tradeoff at play. Any model can be served at a wide range of interactivity levels (tokens/sec per user), and the cost per million tokens (CPMT) shifts accordingly. Mercedes makes metro busses as well as race cars, to follow long with our analogy.
Atomic Claim 333/574 · 2026-02-16_inferencex-v2-nvidia-blackwell-vs::IX2-0350
Atomic Claim 334/574 · 2026-02-16_inferencex-v2-nvidia-blackwell-vs::IX2-0351
Atomic Claim 335/574 · 2026-02-16_inferencex-v2-nvidia-blackwell-vs::IX2-0352
Atomic Claim 336/574 · 2026-02-16_inferencex-v2-nvidia-blackwell-vs::IX2-0353
Atomic Claim 337/574 · 2026-02-16_inferencex-v2-nvidia-blackwell-vs::IX2-0354
Atomic Claim 338/574 · 2026-02-16_inferencex-v2-nvidia-blackwell-vs::IX2-0355
Atomic Claim 339/574 · 2026-02-16_inferencex-v2-nvidia-blackwell-vs::IX2-0356
Atomic Claim 340/574 · 2026-02-16_inferencex-v2-nvidia-blackwell-vs::IX2-0357
Claim: 沿用前述比喻,Mercedes 同時生產大眾巴士與賽車。
Frame:CLAIM_ONLY· Mode:ASSERTED· Mapping:CLAIM_ONLY
開啟逐條審核
Bean counters may think that fast mode is more expensive, but when looking at it through a total cost of ownership lens, fast mode is actually way cheaper for some situations. For example, a GB200 NVL72 rack can cost 3.3 million dollars, and as such, if claude code agentic loops (which runs on Trainium in production) that tool use call NVL72 racks, and these racks run inference 2.5x slower, you would need 2.5x more racks to deliver inference, meaning that not enabling fast mode would cost close to 5 million dollars in extra spend.
Atomic Claim 341/574 · 2026-02-16_inferencex-v2-nvidia-blackwell-vs::IX2-0358
Atomic Claim 342/574 · 2026-02-16_inferencex-v2-nvidia-blackwell-vs::IX2-0359
Atomic Claim 343/574 · 2026-02-16_inferencex-v2-nvidia-blackwell-vs::IX2-0360
Claim: 例如一個 GB200 NVL72 rack 成本約 330 萬美元。
Frame:ATTRIBUTE· Mode:ASSERTED· Mapping:COMPLETE
開啟逐條審核
Atomic Claim 344/574 · 2026-02-16_inferencex-v2-nvidia-blackwell-vs::IX2-0361
Claim: 假設 claude code agentic loops(production 實際運行於 Trainium)需要呼叫 NVL72 racks 執行 tool use。
Frame:ATTRIBUTE· Mode:HYPOTHETICAL· Mapping:COMPLETE
開啟逐條審核
Atomic Claim 345/574 · 2026-02-16_inferencex-v2-nvidia-blackwell-vs::IX2-0362
Claim: 關於 Claude Code:若這些 racks 的 inference 速度慢 2.5 倍,就需要 2.5 倍機櫃才能提供同等 inference capacity,因此不啟用 fast mode 可能需要多支出接近 500 萬美元。
Frame:COMPARISON· Mode:HYPOTHETICAL· Mapping:PARTIAL
開啟逐條審核


Consider a DeepSeek R1 0528 FP4 coding workflow served on B200s with TRT-LLM. At an interactivity of 50 tok/sec/user, inference cost is approximately 4/M output tokens, a 2.5× speed increase for a ~7× price increase, closely mirroring what we see with Anthropic’s fast mode. Note that this assumes DeepSeek R1 is similar to Opus 4.6, which isn’t the case. Still, the general principle holds true.
Atomic Claim 346/574 · 2026-02-16_inferencex-v2-nvidia-blackwell-vs::IX2-0363
Atomic Claim 347/574 · 2026-02-16_inferencex-v2-nvidia-blackwell-vs::IX2-0364
Atomic Claim 348/574 · 2026-02-16_inferencex-v2-nvidia-blackwell-vs::IX2-0365
Atomic Claim 349/574 · 2026-02-16_inferencex-v2-nvidia-blackwell-vs::IX2-0366
Atomic Claim 350/574 · 2026-02-16_inferencex-v2-nvidia-blackwell-vs::IX2-0367

Source: SemiAnalysis InferenceX ↗

Source: SemiAnalysis InferenceX ↗
This follows directly from the fundamental latency-throughput tradeoff in LLM inference. At high batch sizes, GPUs achieve better utilization and greater total token throughput, meaning more users served concurrently and lower cost per token. At low batch sizes with greater parallelism per request, each user gets faster responses, but total token throughput drops. Since the hourly cost of the accelerators ↗ is fixed regardless of how they’re used, lower throughput means fewer tokens over which to amortize that cost, and thus a higher price per token.
Atomic Claim 351/574 · 2026-02-16_inferencex-v2-nvidia-blackwell-vs::IX2-0371
Atomic Claim 352/574 · 2026-02-16_inferencex-v2-nvidia-blackwell-vs::IX2-0368
Atomic Claim 353/574 · 2026-02-16_inferencex-v2-nvidia-blackwell-vs::IX2-0369
Atomic Claim 354/574 · 2026-02-16_inferencex-v2-nvidia-blackwell-vs::IX2-0370
Claim: 在低 batch size、每個 request 使用更多 parallelism 時,每位使用者回應會更快,但總 token throughput 下降。
Frame:ATTRIBUTE· Mode:EXPECTED· Mapping:COMPLETE
開啟逐條審核
In short, fast mode isn’t necessarily a hardware story, but merely the natural consequence of trading throughput for latency on the same GPUs.
Atomic Claim 355/574 · 2026-02-16_inferencex-v2-nvidia-blackwell-vs::IX2-0372

Source: SemiAnalysis InferenceX ↗
Furthermore, we observe that inference optimization techniques such as speculative decoding, as explained earlier, can directly lead to cheaper inference; no new chips are required.
Atomic Claim 356/574 · 2026-02-16_inferencex-v2-nvidia-blackwell-vs::IX2-0373
Claim: 此外,前文所述的 speculative decoding 等 inference optimization techniques,可直接降低 inference 成本。
Frame:ATTRIBUTE· Mode:ASSERTED· Mapping:COMPLETE
開啟逐條審核
Atomic Claim 357/574 · 2026-02-16_inferencex-v2-nvidia-blackwell-vs::IX2-0374
Claim: 關於 speculative decoding:這些改善不需要新的晶片。
Frame:ATTRIBUTE· Mode:ASSERTED· Mapping:PARTIAL
開啟逐條審核
Take the following example, DeepSeek R1 FP4 on an 8k/1k workload. At an interactivity level of 150 tok/sec/user, the baseline GB300 Dynamo TRT cost per million tokens is approximately 0.11. This is a ~21x price decrease at this interactivity level simply by employing an inference optimization technique.
Atomic Claim 358/574 · 2026-02-16_inferencex-v2-nvidia-blackwell-vs::IX2-0375
Atomic Claim 359/574 · 2026-02-16_inferencex-v2-nvidia-blackwell-vs::IX2-0376
Atomic Claim 360/574 · 2026-02-16_inferencex-v2-nvidia-blackwell-vs::IX2-0377
Atomic Claim 361/574 · 2026-02-16_inferencex-v2-nvidia-blackwell-vs::IX2-0378

Source: SemiAnalysis InferenceX ↗

Source: SemiAnalysis InferenceX ↗

Source: SemiAnalysis InferenceX ↗
Fixing an interactivity level of 50 tok/sec/user, we further see how much MTP can effectively decrease CPMT across a variety of chips.
Atomic Claim 362/574 · 2026-02-16_inferencex-v2-nvidia-blackwell-vs::IX2-0379

Source: SemiAnalysis InferenceX ↗
Wide Expert Parallelism (WideEP) and Disaggregated Prefill
In this section, we will go deeper on expert parallelism and go on to explain what _wide _expert parallelism is. We will then explain the idea of Disaggregated Prefill, how it is different from WideEP, and how WideEP and Disaggregated Prefill are used in unison to achieve SOTA performance.
Atomic Claim 363/574 · 2026-02-16_inferencex-v2-nvidia-blackwell-vs::IX2-0380
Claim: 本節會更深入討論 expert parallelism,並說明 wide expert parallelism 的概念。
Frame:ATTRIBUTE· Mode:ASSERTED· Mapping:PARTIAL
開啟逐條審核
Atomic Claim 364/574 · 2026-02-16_inferencex-v2-nvidia-blackwell-vs::IX2-0381
Claim: 接著會介紹 Disaggregated Prefill,以及它與 WideEP 的差異。
Frame:ATTRIBUTE· Mode:ASSERTED· Mapping:PARTIAL
開啟逐條審核
Atomic Claim 365/574 · 2026-02-16_inferencex-v2-nvidia-blackwell-vs::IX2-0382
Claim: 並說明 WideEP 與 Disaggregated Prefill 如何搭配使用以達成 SOTA performance。
Frame:ATTRIBUTE· Mode:ASSERTED· Mapping:PARTIAL
開啟逐條審核
WideEP
By now, most frontier AI labs employ Mixture of Experts (MoE) model architectures as opposed to dense. In MoE architectures, only a subset of “experts” are activated for each token. For instance, DeepSeek R1 has 671B total parameters, but only 37B active parameters. Specifically, DeepSeek R1 has 256 routed experts (and 1 shared expert) with each token being routed to 8 distinct experts. This architecture lends itself naturally to expert parallelism (EP), which evenly distributes expert weights across some number of GPUs.
Atomic Claim 366/574 · 2026-02-16_inferencex-v2-nvidia-blackwell-vs::IX2-0383
Claim: 目前大多數 frontier AI labs 都採用 Mixture of Experts (MoE) architecture,而不是 dense model。
Frame:ATTRIBUTE· Mode:ASSERTED· Mapping:COMPLETE
開啟逐條審核
Atomic Claim 367/574 · 2026-02-16_inferencex-v2-nvidia-blackwell-vs::IX2-0384
Atomic Claim 368/574 · 2026-02-16_inferencex-v2-nvidia-blackwell-vs::IX2-0385
Atomic Claim 369/574 · 2026-02-16_inferencex-v2-nvidia-blackwell-vs::IX2-0386
Atomic Claim 370/574 · 2026-02-16_inferencex-v2-nvidia-blackwell-vs::IX2-0387
Claim: 這種架構天然適合 expert parallelism (EP),也就是把 expert weights 平均分配到多顆 GPUs。
Frame:ATTRIBUTE· Mode:ASSERTED· Mapping:COMPLETE
開啟逐條審核
Consider serving DeepSeek R1 on a single 8-GPU server. At 671B parameters, some form of parallelism is required to fit the model across available HBM. The naive approach is tensor parallelism (TP), which shards every weight matrix across all GPUs. This works well for dense models but ignores the sparse activation pattern of MoE. With TP=8, each expert’s weights are sharded across all 8 GPUs, meaning every expert activation requires an all-reduce across all GPUs & the reduction dims of the GEMM is smaller leading to lower arithmetic intensity, even though only 8 of 256 experts activate per token. TP treats each expert like a dense layer, paying full cross-GPU communication cost while the model’s sparsity goes unexploited.
Atomic Claim 371/574 · 2026-02-16_inferencex-v2-nvidia-blackwell-vs::IX2-0388
Atomic Claim 372/574 · 2026-02-16_inferencex-v2-nvidia-blackwell-vs::IX2-0389
Atomic Claim 373/574 · 2026-02-16_inferencex-v2-nvidia-blackwell-vs::IX2-0390
Claim: 最直接的方法是 tensor parallelism (TP),把每個 weight matrix 分片到所有 GPUs。
Frame:RELATION· Mode:ASSERTED· Mapping:COMPLETE
開啟逐條審核
Atomic Claim 374/574 · 2026-02-16_inferencex-v2-nvidia-blackwell-vs::IX2-0391
Atomic Claim 375/574 · 2026-02-16_inferencex-v2-nvidia-blackwell-vs::IX2-0392
Claim: 在 TP=8 下,每個 expert weights 都分散在 8 顆 GPUs,因此每次 expert activation 都需要跨所有 GPUs 執行 all-reduce;同時 GEMM reduction dimension 變小,降低 arithmetic intensity,即使每 token 實際只啟動 256 個 experts 中的 8 個。
Frame:ATTRIBUTE· Mode:ASSERTED· Mapping:COMPLETE
開啟逐條審核
Atomic Claim 376/574 · 2026-02-16_inferencex-v2-nvidia-blackwell-vs::IX2-0393
Expert parallelism takes a more well-suited approach, assigning whole experts to individual GPUs. With EP=8, we divide the 256 experts per layer across 8 GPUs for a total of 32 experts/layer/GPU. Each GPU holds approximately 1/8th of the expert weights plus a full replica of the non-expert weights (attention projections, embeddings, normalization, and the shared expert). Since roughly 90%+ of DeepSeek R1’s parameters are routed expert weights, EP captures most of the memory savings, and replicating the remaining less than 30B non-expert parameters across all 8 GPUs is affordable.
Atomic Claim 377/574 · 2026-02-16_inferencex-v2-nvidia-blackwell-vs::IX2-0394
Claim: Expert parallelism 更適合這種架構,它直接把完整 expert 分配給個別 GPUs。
Frame:ATTRIBUTE· Mode:ASSERTED· Mapping:COMPLETE
開啟逐條審核
Atomic Claim 378/574 · 2026-02-16_inferencex-v2-nvidia-blackwell-vs::IX2-0395
Atomic Claim 379/574 · 2026-02-16_inferencex-v2-nvidia-blackwell-vs::IX2-0396
Atomic Claim 380/574 · 2026-02-16_inferencex-v2-nvidia-blackwell-vs::IX2-0397
Atomic Claim 381/574 · 2026-02-16_inferencex-v2-nvidia-blackwell-vs::IX2-0398
The forward pass proceeds in two phases per layer. During attention, each GPU acts as an independent data-parallel rank, processing its own subset of requests using its replicated non-expert weights, no inter-GPU communication is needed. During the MoE phase, a lightweight router determines which experts each token requires, and tokens are dispatched to the appropriate GPUs via all-to-all communication. Each GPU executes its local experts on only the tokens routed to it, and results are returned via a second all-to-all.
Atomic Claim 382/574 · 2026-02-16_inferencex-v2-nvidia-blackwell-vs::IX2-0399
Atomic Claim 383/574 · 2026-02-16_inferencex-v2-nvidia-blackwell-vs::IX2-0400
Atomic Claim 384/574 · 2026-02-16_inferencex-v2-nvidia-blackwell-vs::IX2-0401
Atomic Claim 385/574 · 2026-02-16_inferencex-v2-nvidia-blackwell-vs::IX2-0402
Atomic Claim 386/574 · 2026-02-16_inferencex-v2-nvidia-blackwell-vs::IX2-0403
Atomic Claim 387/574 · 2026-02-16_inferencex-v2-nvidia-blackwell-vs::IX2-0404

_An EP8 DP8 deployment of DeepSeek R1. All 256 experts per layer are divided evenly among the 8 GPUs, whereas attention along with other non-expert weights (shared expert, gating network, RMSNorm, LM head, etc.) are replicated across all 8 DP ranks. _Source: SemiAnalysis
Atomic Claim 388/574 · 2026-02-16_inferencex-v2-nvidia-blackwell-vs::IX2-0405
Atomic Claim 389/574 · 2026-02-16_inferencex-v2-nvidia-blackwell-vs::IX2-0406
The obvious way to scale is replication: deploy N independent EP8 instances across N nodes. Each instance serves requests independently with no cross-node communication. This scales throughput linearly, but each GPU still holds 32 experts per layer, and each token activates at most 8 of those 32 local experts. 75% of expert weights sit cold in HBM.
Atomic Claim 390/574 · 2026-02-16_inferencex-v2-nvidia-blackwell-vs::IX2-0408
Atomic Claim 391/574 · 2026-02-16_inferencex-v2-nvidia-blackwell-vs::IX2-0409
Atomic Claim 392/574 · 2026-02-16_inferencex-v2-nvidia-blackwell-vs::IX2-0410
Atomic Claim 393/574 · 2026-02-16_inferencex-v2-nvidia-blackwell-vs::IX2-0411
Wide expert parallelism (WideEP) takes a different approach by scaling EP _across _nodes rather than replicating independent instances. On a 64-GPU cluster (8 nodes), DP64/EP64 places only 256/64 = 4 experts per layer per GPU, each still holding a full replica of the non-expert weights. During the MoE phase, tokens from all 64 DP ranks are dispatched via all-to-all to the GPUs hosting their routed experts.
Atomic Claim 394/574 · 2026-02-16_inferencex-v2-nvidia-blackwell-vs::IX2-0412
Claim: Wide expert parallelism(WideEP)採用不同方式:不是複製獨立 instances,而是把 EP 直接跨 nodes 擴展。
Frame:ATTRIBUTE· Mode:ASSERTED· Mapping:COMPLETE
開啟逐條審核
Atomic Claim 395/574 · 2026-02-16_inferencex-v2-nvidia-blackwell-vs::IX2-0413
Atomic Claim 396/574 · 2026-02-16_inferencex-v2-nvidia-blackwell-vs::IX2-0414
This yields three compounding benefits over the single-node EP8 baseline. First, reducing expert footprint from 32 to 4 experts/GPU frees substantial HBM for KV cache, directly increasing per-GPU batch size capacity. Second, 64 DP ranks funneling tokens through fewer experts per GPU increases tokens-per-expert, raising arithmetic intensity (more FLOPs per byte of weights loaded) and improving compute utilization. The same expert weights service 8x more tokens per step. Third, aggregate HBM bandwidth scales linearly with GPU count; 64 GPUs loading expert weights simultaneously provide 8x the memory bandwidth of a single node, reducing memory bottleneck.
Atomic Claim 397/574 · 2026-02-16_inferencex-v2-nvidia-blackwell-vs::IX2-0415
Atomic Claim 398/574 · 2026-02-16_inferencex-v2-nvidia-blackwell-vs::IX2-0416
Atomic Claim 399/574 · 2026-02-16_inferencex-v2-nvidia-blackwell-vs::IX2-0417
Claim: 第二,64 個 DP ranks 的 tokens 被集中到每顆 GPU 較少的 experts 上,使每個 expert 接收到更多 tokens,提高 arithmetic intensity,也就是每載入一 byte weights 可執行更多 FLOPs,進而提升 compute utilization。
Frame:COMPARISON· Mode:ASSERTED· Mapping:COMPLETE
開啟逐條審核
Atomic Claim 400/574 · 2026-02-16_inferencex-v2-nvidia-blackwell-vs::IX2-0418
Atomic Claim 401/574 · 2026-02-16_inferencex-v2-nvidia-blackwell-vs::IX2-0419
Atomic Claim 402/574 · 2026-02-16_inferencex-v2-nvidia-blackwell-vs::IX2-0420

_A WideEP EP64 DP64 deployment of DeepSeek R1. All 256 experts per layer are divided evenly among the 64 GPUs (8 nodes), and attention and other non-expert weights (shared expert, gating network, RMSNorm, LM head, etc.) are replicated across all 64 DP ranks. _Source: SemiAnalysis
Atomic Claim 403/574 · 2026-02-16_inferencex-v2-nvidia-blackwell-vs::IX2-0421
Atomic Claim 404/574 · 2026-02-16_inferencex-v2-nvidia-blackwell-vs::IX2-0422
The above configurations use only DP+EP (also known as DEP), where each GPU holds a full replica of all non-expert weights. As GPU count grows, this replication becomes increasingly wasteful. On a 64-GPU DP64/EP64 deployment, every GPU stores an identical copy of the ~40B non-expert parameters.
Atomic Claim 405/574 · 2026-02-16_inferencex-v2-nvidia-blackwell-vs::IX2-0424
Atomic Claim 406/574 · 2026-02-16_inferencex-v2-nvidia-blackwell-vs::IX2-0425
Atomic Claim 407/574 · 2026-02-16_inferencex-v2-nvidia-blackwell-vs::IX2-0426
Adding tensor parallelism within groups of GPUs addresses this. In an EP64/DP8/TP8 configuration, the 64 GPUs are organized into 8 DP groups of 8 GPUs each. Within each TP group, the attention projections, shared expert, normalization, and LM head are sharded 8 ways, so each GPU holds only 1/8th of the non-expert weights. Across the full cluster, the 256 experts are still distributed one-per-4-GPUs as before.
Atomic Claim 408/574 · 2026-02-16_inferencex-v2-nvidia-blackwell-vs::IX2-0427
Claim: 可在 GPUs 子群組內加入 tensor parallelism 解決這個問題。
Frame:ATTRIBUTE· Mode:ASSERTED· Mapping:COMPLETE
開啟逐條審核
Atomic Claim 409/574 · 2026-02-16_inferencex-v2-nvidia-blackwell-vs::IX2-0428
Atomic Claim 410/574 · 2026-02-16_inferencex-v2-nvidia-blackwell-vs::IX2-0429
Atomic Claim 411/574 · 2026-02-16_inferencex-v2-nvidia-blackwell-vs::IX2-0430
Atomic Claim 412/574 · 2026-02-16_inferencex-v2-nvidia-blackwell-vs::IX2-0431
Pure DEP has a single communication pattern: all-to-all for expert routing. Adding TP introduces a second all-reduce within each TP group for the attention and non-expert computations. The key design principle is to place TP groups within a single node, where NVLink or MNNVL provides high-bandwidth interconnect, and run EP/DP across nodes, where the all-to-all communication pattern can tolerate higher latency.
Atomic Claim 413/574 · 2026-02-16_inferencex-v2-nvidia-blackwell-vs::IX2-0432
Atomic Claim 414/574 · 2026-02-16_inferencex-v2-nvidia-blackwell-vs::IX2-0433
Claim: 加入 TP 後,attention 與 non-expert computations 會在每個 TP group 內多出第二種 all-reduce 通訊。
Frame:ATTRIBUTE· Mode:ASSERTED· Mapping:COMPLETE
開啟逐條審核
Atomic Claim 415/574 · 2026-02-16_inferencex-v2-nvidia-blackwell-vs::IX2-0434
Atomic Claim 416/574 · 2026-02-16_inferencex-v2-nvidia-blackwell-vs::IX2-0435
As always, the tradeoff is that of throughput versus latency. TP=8 within a group means those 8 GPUs now share a batch and must synchronize every decode step, reducing effective DP degree from 64 to 8. Per-GPU batching independence on the attention side is lost. But each DP group now processes attention 8x faster per step, since the matmul is split 8 ways across the TP group. Per-token latency drops while peak concurrency also drops, sliding the configuration along the latency-throughput Pareto frontier relative to pure DEP.
Atomic Claim 417/574 · 2026-02-16_inferencex-v2-nvidia-blackwell-vs::IX2-0436
Atomic Claim 418/574 · 2026-02-16_inferencex-v2-nvidia-blackwell-vs::IX2-0437
Atomic Claim 419/574 · 2026-02-16_inferencex-v2-nvidia-blackwell-vs::IX2-0438
Atomic Claim 420/574 · 2026-02-16_inferencex-v2-nvidia-blackwell-vs::IX2-0439
Atomic Claim 421/574 · 2026-02-16_inferencex-v2-nvidia-blackwell-vs::IX2-0440
Disaggregated Prefill
Disaggregated prefill, sometimes referred to as prefill-decode (PD) disaggregation, is the process of performing prefill and decode phases of LLM inference on separate nodes. Prefill occurs when a request is first processed, and a forward pass is computed on all tokens at once, thereby “prefilling” the KV cache for this request. This is a compute-intensive operation as all tokens feed through the forward pass in parallel. Tokens are then generated or “decoded” one at a time, loading the KV cache from HBM at each decode step. This is a memory-intensive process as the growing KV cache is constantly being loaded.
Atomic Claim 422/574 · 2026-02-16_inferencex-v2-nvidia-blackwell-vs::IX2-0441
Atomic Claim 423/574 · 2026-02-16_inferencex-v2-nvidia-blackwell-vs::IX2-0442
Atomic Claim 424/574 · 2026-02-16_inferencex-v2-nvidia-blackwell-vs::IX2-0443
Atomic Claim 425/574 · 2026-02-16_inferencex-v2-nvidia-blackwell-vs::IX2-0444
Claim: 關於 disaggregated prefill:因為所有 tokens 都同時通過 forward pass,因此這是 compute-intensive operation。
Frame:ATTRIBUTE· Mode:ASSERTED· Mapping:COMPLETE
開啟逐條審核
Atomic Claim 426/574 · 2026-02-16_inferencex-v2-nvidia-blackwell-vs::IX2-0445
Atomic Claim 427/574 · 2026-02-16_inferencex-v2-nvidia-blackwell-vs::IX2-0446
In traditional single-node inference, engines interleave prefill and decode on the same GPUs. Incoming prefill requests stall in-flight decode batches, increasing both time-to-first-token and inter-token latency. Chunked prefill mitigates this by breaking long prefills into smaller pieces, but the fundamental resource contention remains. Disaggregated prefill eliminates this entirely!
Atomic Claim 428/574 · 2026-02-16_inferencex-v2-nvidia-blackwell-vs::IX2-0447
Atomic Claim 429/574 · 2026-02-16_inferencex-v2-nvidia-blackwell-vs::IX2-0448
Atomic Claim 430/574 · 2026-02-16_inferencex-v2-nvidia-blackwell-vs::IX2-0449
Atomic Claim 431/574 · 2026-02-16_inferencex-v2-nvidia-blackwell-vs::IX2-0450
Atomic Claim 432/574 · 2026-02-16_inferencex-v2-nvidia-blackwell-vs::IX2-0451
Claim: Disaggregated prefill 可以把這種 contention 完全消除。
Frame:ATTRIBUTE· Mode:ASSERTED· Mapping:PARTIAL
開啟逐條審核

Source: DistServe ↗
Disaggregation also enables independent scaling and optimization of each phase. With separate nodes, each phase can be tuned independently: different parallelism strategies, different batch sizes, and different memory allocation ratios. The ratio of prefill to decode nodes can also be matched to the workload’s input-output length ratio. For instance, prefill-dominated workloads (long input, short output e.g., summarization, RAG, agentic coding with large context windows) allocate more prefill instances. Decode-dominated workloads (short input, long output e.g., chain-of-thought reasoning, long-form generation) allocate more decode instances. Workloads with high cache hit rates also tend toward more decode, since reused KV cache entries from shared system prompts or multi-turn conversation history skip prefill entirely.
Atomic Claim 433/574 · 2026-02-16_inferencex-v2-nvidia-blackwell-vs::IX2-0452
Claim: 關於 disaggregated prefill:Disaggregation 也讓兩個階段可以各自獨立 scale 與 optimize。
Frame:ATTRIBUTE· Mode:ASSERTED· Mapping:COMPLETE
開啟逐條審核
Atomic Claim 434/574 · 2026-02-16_inferencex-v2-nvidia-blackwell-vs::IX2-0453
Claim: 在 disaggregated serving 中,prefill 與 decode 可以採用不同 parallelism strategies。
Frame:NARY_RELATION· Mode:ASSERTED· Mapping:COMPLETE
開啟逐條審核
Atomic Claim 435/574 · 2026-02-16_inferencex-v2-nvidia-blackwell-vs::IX2-0454
Claim: 在 disaggregated serving 中,prefill 與 decode 可以使用不同 batch sizes。
Frame:NARY_RELATION· Mode:ASSERTED· Mapping:COMPLETE
開啟逐條審核
Atomic Claim 436/574 · 2026-02-16_inferencex-v2-nvidia-blackwell-vs::IX2-0455
Claim: 在 disaggregated serving 中,prefill 與 decode 可以使用不同 memory-allocation ratios。
Frame:NARY_RELATION· Mode:ASSERTED· Mapping:COMPLETE
開啟逐條審核
Atomic Claim 437/574 · 2026-02-16_inferencex-v2-nvidia-blackwell-vs::IX2-0456
Atomic Claim 438/574 · 2026-02-16_inferencex-v2-nvidia-blackwell-vs::IX2-0457
Claim: 例如 prefill-dominated workloads,也就是長 input、短 output 的 summarization、RAG、具大型 context windows 的 agentic coding,會配置更多 prefill instances。
Frame:RELATION· Mode:ASSERTED· Mapping:COMPLETE
開啟逐條審核
Atomic Claim 439/574 · 2026-02-16_inferencex-v2-nvidia-blackwell-vs::IX2-0458
Atomic Claim 440/574 · 2026-02-16_inferencex-v2-nvidia-blackwell-vs::IX2-0459
The key cost of disaggregation is KV cache transfer. After prefill completes, the full KV cache for that request must be transmitted from the prefill node to the decode node before the first decode token can be generated. For a model like DeepSeek R1 with 61 layers and FP8 KV cache, an 8192-token prefill produces roughly 500MB of KV data that must cross the network, adding directly to TTFT. This transfer is performed over RDMA (typically RoCE or InfiniBand) using zero-copy GPU-to-GPU data movement without CPU involvement. Libraries like NIXL (NVIDIA Inference Transfer Library) abstract the data movement layer behind a unified asynchronous API with pluggable backends for UCX, GPUDirect Storage, and other transports. This decouples the inference engine from any specific transfer protocol and enables disaggregation across heterogeneous hardware where prefill and decode instances may span different device types or interconnects.
Atomic Claim 441/574 · 2026-02-16_inferencex-v2-nvidia-blackwell-vs::IX2-0460
Atomic Claim 442/574 · 2026-02-16_inferencex-v2-nvidia-blackwell-vs::IX2-0461
Atomic Claim 443/574 · 2026-02-16_inferencex-v2-nvidia-blackwell-vs::IX2-0462
Atomic Claim 444/574 · 2026-02-16_inferencex-v2-nvidia-blackwell-vs::IX2-0463
Atomic Claim 445/574 · 2026-02-16_inferencex-v2-nvidia-blackwell-vs::IX2-0464
Claim: 像 NIXL(NVIDIA Inference Transfer Library)這類 library,透過統一 asynchronous API 抽象化 data-movement layer,並提供 UCX、GPUDirect Storage 與其他 transports 的 pluggable backends。
Frame:NARY_RELATION· Mode:ASSERTED· Mapping:PARTIAL
開啟逐條審核
Atomic Claim 446/574 · 2026-02-16_inferencex-v2-nvidia-blackwell-vs::IX2-0465
Claim: 這使 inference engine 不必綁定特定 transfer protocol,也能支援 heterogeneous hardware 的 disaggregation,讓 prefill 與 decode instances 跨不同 device types 或 interconnects。
Frame:NARY_RELATION· Mode:ASSERTED· Mapping:COMPLETE
開啟逐條審核

Optimizing Inference with Wide EP + Disaggregated Serving
Wide EP and disaggregated prefill are separate techniques that are often used together to achieve Pareto optimal performance. In this section, we walk through real results from InferenceX to build intuition for which combinations of parallelism strategy, wide EP, and disaggregated prefill are appropriate at different interactivity levels.
Atomic Claim 447/574 · 2026-02-16_inferencex-v2-nvidia-blackwell-vs::IX2-0466
Claim: Wide EP 與 disaggregated prefill 是兩種獨立技術,但經常搭配使用 together,以達到 Pareto-optimal performance。
Frame:ATTRIBUTE· Mode:ASSERTED· Mapping:PARTIAL
開啟逐條審核
Atomic Claim 448/574 · 2026-02-16_inferencex-v2-nvidia-blackwell-vs::IX2-0467
Claim: 本節透過真實 InferenceX 結果,建立在不同 interactivity 水準下如何選擇 parallelism strategy 與 wide EP 的直覺。
Frame:ATTRIBUTE· Mode:ASSERTED· Mapping:COMPLETE
開啟逐條審核
Atomic Claim 449/574 · 2026-02-16_inferencex-v2-nvidia-blackwell-vs::IX2-0468
Claim: 並進一步判斷何時適合使用 disaggregated prefill。
Frame:ATTRIBUTE· Mode:ASSERTED· Mapping:PARTIAL
開啟逐條審核
It helps to first understand what parallelism strategies fall on what parts of the Pareto frontier for single-node configurations. Take the example of DeepSeek R1 FP4 8k/1k on a single 8-GPU B200 node with TRT-LLM. The optimal strategy shifts as you move along the frontier, driven primarily by batch size and its effect on expert activation density.
Atomic Claim 450/574 · 2026-02-16_inferencex-v2-nvidia-blackwell-vs::IX2-0470
Atomic Claim 451/574 · 2026-02-16_inferencex-v2-nvidia-blackwell-vs::IX2-0471
Claim: 沿著 frontier 移動時,最佳 strategy 會改變,主要由 batch size 以及它對 expert activation density 的影響所決定。
Frame:ATTRIBUTE· Mode:ASSERTED· Mapping:PARTIAL
開啟逐條審核
At the highest interactivity levels (batch 1-16), pure TP outperforms any configuration involving EP. At low batch sizes, only a small fraction of experts activate per step. With EP, these activations are distributed unevenly across GPUs: at batch 4, only 32 of 256 experts fire, and any given GPU has roughly a low double digit percent chance of receiving zero routed tokens in a given layer. TP avoids this by sharding every expert across all GPUs, so all 8 GPUs participate equally in every expert computation regardless of which experts the router selects. We collected expert activation ratio versus batch size data while profiling DeepSeek R1, which confirms that at batch sizes 16 and below, expert activation per layer is very low.
Atomic Claim 452/574 · 2026-02-16_inferencex-v2-nvidia-blackwell-vs::IX2-0472
Atomic Claim 453/574 · 2026-02-16_inferencex-v2-nvidia-blackwell-vs::IX2-0473
Claim: 在低 batch size 下,每個 step 只會啟動少部分 experts。
Frame:ATTRIBUTE· Mode:ASSERTED· Mapping:COMPLETE
開啟逐條審核
Atomic Claim 454/574 · 2026-02-16_inferencex-v2-nvidia-blackwell-vs::IX2-0474
Atomic Claim 455/574 · 2026-02-16_inferencex-v2-nvidia-blackwell-vs::IX2-0475
Atomic Claim 456/574 · 2026-02-16_inferencex-v2-nvidia-blackwell-vs::IX2-0476
Atomic Claim 457/574 · 2026-02-16_inferencex-v2-nvidia-blackwell-vs::IX2-0477
Claim: InferenceX 在 profiling DeepSeek R1 時收集 expert activation ratio vs batch size,結果確認 batch 16 以下每層 expert activation 都很低。
Frame:COMPARISON· Mode:ASSERTED· Mapping:PARTIAL
開啟逐條審核

Source: SemiAnalysis
As we move to slightly lower interactivities, batch sizes remain small enough that expert weights are still sharded via TP rather than EP. The crossover occurs around batch 32, where approximately 50-60% of experts activate per layer. At this density, EP’s load imbalance becomes tolerable and its token-routing overhead is cheaper than the per-expert all-reduce required by TP. Configurations in this range use TEP: tensor parallelism for attention (all GPUs collaborate on each attention computation), expert parallelism for MoE layers (experts assigned to specific GPUs with all-to-all routing). In the highest throughput, lowest interactivity region of the frontier, batch sizes are large (128+) and configurations shift to full DEP: attention weights are fully replicated across all GPUs as independent data-parallel ranks, experts are distributed via EP, and batch capacity is maximized at the cost of per-token latency. (128+) and attention weights are fully replicated across all DP ranks, maximizing throughput.
Atomic Claim 458/574 · 2026-02-16_inferencex-v2-nvidia-blackwell-vs::IX2-0478
Claim: 當 interactivity 稍微下降時,batch size 仍偏小,因此 expert weights 仍較適合用 TP 分片,而非 EP。
Frame:RELATION· Mode:ASSERTED· Mapping:COMPLETE
開啟逐條審核
Atomic Claim 459/574 · 2026-02-16_inferencex-v2-nvidia-blackwell-vs::IX2-0480
Claim: 到這個密度後,EP 的 load imbalance 已可接受,而 token-routing overhead 也低於 TP 對每個 expert 都需執行的 all-reduce 成本。
Frame:COMPARISON· Mode:ASSERTED· Mapping:PARTIAL
開啟逐條審核
Atomic Claim 460/574 · 2026-02-16_inferencex-v2-nvidia-blackwell-vs::IX2-0481
Claim: 此區間 configurations 使用 TEP:attention 採 tensor parallelism,由所有 GPUs 協同完成每次 attention;MoE layers 則採 expert parallelism,把 experts 分配到特定 GPUs 並透過 all-to-all routing。
Frame:NARY_RELATION· Mode:ASSERTED· Mapping:COMPLETE
開啟逐條審核
Atomic Claim 461/574 · 2026-02-16_inferencex-v2-nvidia-blackwell-vs::IX2-0482

Source: SemiAnalysis InferenceX ↗
We observe the same general pattern when extending to wide EP with disaggregated prefill. Prefill and decode run with separate parallelism strategies and node counts, both tuned to the workload and target interactivity level. Take an 8k/1k workload (prefill heavy) at the high-throughput, low-interactivity end of the frontier. Prefill is the bottleneck as each request requires a forward pass of 8192 input tokens, which is computationally expensive. Recipes in this region allocate more prefill nodes than decode (4P1D, 7P2D, 4P3D) to sustain high prefill throughput. These prefill nodes run DEP configurations, replicating attention weights across independent data-parallel ranks so that multiple long-context prefills can be processed simultaneously. Decode nodes are fewer but run wide DEP with large batch sizes by the same principle as with single node.
Atomic Claim 462/574 · 2026-02-16_inferencex-v2-nvidia-blackwell-vs::IX2-0485
Claim: 把範圍擴展到搭配 disaggregated prefill 的 wide EP 時,也可觀察到相同一般模式。
Frame:COMPARISON· Mode:ASSERTED· Mapping:COMPLETE
開啟逐條審核
Atomic Claim 463/574 · 2026-02-16_inferencex-v2-nvidia-blackwell-vs::IX2-0486
Atomic Claim 464/574 · 2026-02-16_inferencex-v2-nvidia-blackwell-vs::IX2-0487
Atomic Claim 465/574 · 2026-02-16_inferencex-v2-nvidia-blackwell-vs::IX2-0488
Atomic Claim 466/574 · 2026-02-16_inferencex-v2-nvidia-blackwell-vs::IX2-0489
Atomic Claim 467/574 · 2026-02-16_inferencex-v2-nvidia-blackwell-vs::IX2-0490
Atomic Claim 468/574 · 2026-02-16_inferencex-v2-nvidia-blackwell-vs::IX2-0491
Atomic Claim 469/574 · 2026-02-16_inferencex-v2-nvidia-blackwell-vs::IX2-0492
Atomic Claim 470/574 · 2026-02-16_inferencex-v2-nvidia-blackwell-vs::IX2-0493
On the low interactivity end of the frontier, there are fewer concurrent requests in flight, so a single prefill instance can keep pace with incoming demand. Yet each request still requires 1024 decode steps, and at high interactivity those steps must be fast. Recipes in this region shift to more decode nodes than prefill (1P3D, 1P4D), with each decode instance running TEP at low batch size. Tensor parallelism on attention minimizes per-step latency by sharding the computation across all GPUs in the instance, while expert parallelism handles MoE routing at the moderate batch sizes where EP load balance is sufficient. Multiple small-batch decode instances, rather than fewer large-batch ones, keep per-token latency low while still providing enough concurrent serving capacity.
Atomic Claim 471/574 · 2026-02-16_inferencex-v2-nvidia-blackwell-vs::IX2-0494
Atomic Claim 472/574 · 2026-02-16_inferencex-v2-nvidia-blackwell-vs::IX2-0495
Atomic Claim 473/574 · 2026-02-16_inferencex-v2-nvidia-blackwell-vs::IX2-0496
Atomic Claim 474/574 · 2026-02-16_inferencex-v2-nvidia-blackwell-vs::IX2-0497
Atomic Claim 475/574 · 2026-02-16_inferencex-v2-nvidia-blackwell-vs::IX2-0498
Claim: 每個 decode instance 在低 batch size 下執行 TEP。
Frame:ATTRIBUTE· Mode:ASSERTED· Mapping:COMPLETE
開啟逐條審核
Atomic Claim 476/574 · 2026-02-16_inferencex-v2-nvidia-blackwell-vs::IX2-0499
Claim: Attention 採 Tensor parallelism,把運算分散到 instance 中所有 GPUs,以降低每 step latency。
Frame:ATTRIBUTE· Mode:ASSERTED· Mapping:COMPLETE
開啟逐條審核
Atomic Claim 477/574 · 2026-02-16_inferencex-v2-nvidia-blackwell-vs::IX2-0500
Claim: expert parallelism 則在中等 batch size、EP load balance 已足夠時負責 MoE routing。
Frame:RELATION· Mode:ASSERTED· Mapping:COMPLETE
開啟逐條審核
Atomic Claim 478/574 · 2026-02-16_inferencex-v2-nvidia-blackwell-vs::IX2-0501



Source: SemiAnalysis InferenceX
Dive into DeepSeek R1 Single Node Results
On DeepSeek R1 FP8 1k1k, we see that MI355X is competitive with its counterpart B200 on single node scenarios, despite getting mogged on FP4 multi node scenarios. MI355X (SGLang) even beats B200 (SGLang) in throughput performance at lower interactivity levels. Moreover, MI355X (SGLang) beats B200 (TRT and SGLang) in most cases from a perf/TCO perspective.
Atomic Claim 479/574 · 2026-02-16_inferencex-v2-nvidia-blackwell-vs::IX2-0502
Atomic Claim 480/574 · 2026-02-16_inferencex-v2-nvidia-blackwell-vs::IX2-0503
Atomic Claim 481/574 · 2026-02-16_inferencex-v2-nvidia-blackwell-vs::IX2-0504
Unfortunately, the year is 2026, and most frontier labs and inference providers are not running FP8 nor single node inference.
Atomic Claim 482/574 · 2026-02-16_inferencex-v2-nvidia-blackwell-vs::IX2-0505
Claim: 然而,現在已經是 2026 年。
Frame:CLAIM_ONLY· Mode:ASSERTED· Mapping:CLAIM_ONLY
開啟逐條審核
Atomic Claim 483/574 · 2026-02-16_inferencex-v2-nvidia-blackwell-vs::IX2-0506
This result goes to show that AMDs chips are great and can be extremely competitive with Nvidia if only they could move faster on the software front. Speed is the moat.
Atomic Claim 484/574 · 2026-02-16_inferencex-v2-nvidia-blackwell-vs::IX2-0507
Atomic Claim 485/574 · 2026-02-16_inferencex-v2-nvidia-blackwell-vs::IX2-0508
Atomic Claim 486/574 · 2026-02-16_inferencex-v2-nvidia-blackwell-vs::IX2-0509
Claim: 關於 DeepSeek R1:速度就是護城河。
Frame:ATTRIBUTE· Mode:ASSERTED· Mapping:PARTIAL
開啟逐條審核

Source: SemiAnalysis InferenceX ↗

Source: SemiAnalysis InferenceMAX ↗
To that end, we see MI355X fall well behind B200 in performance on FP4:
Atomic Claim 487/574 · 2026-02-16_inferencex-v2-nvidia-blackwell-vs::IX2-0510

Source: SemiAnalysis InferenceX ↗
In comparing DeepSeek R1 FP8 perf between H200 (SGLang) and MI325X (SGLang), not much has changed since our initial release of InferenceXv1 last October. The MI325X data was captured on Feb 12th, 2026 with SGLang 0.5.8 whereas the B200 data was captured Jan 23, 2026 with SGLang 0.5.7.
Atomic Claim 488/574 · 2026-02-16_inferencex-v2-nvidia-blackwell-vs::IX2-0511
Atomic Claim 489/574 · 2026-02-16_inferencex-v2-nvidia-blackwell-vs::IX2-0512
One thing we note is the considerably smaller interactivity range for MI325X than H200, with H200 ranging from 30-90 tok/sec/user whereas MI325X ranges from only 13-35 tok/sec/user. This is problematic for providers who would like to serve users at a broader range of interactivity.
Atomic Claim 490/574 · 2026-02-16_inferencex-v2-nvidia-blackwell-vs::IX2-0513
Atomic Claim 491/574 · 2026-02-16_inferencex-v2-nvidia-blackwell-vs::IX2-0514
Claim: 關於 DeepSeek R1:SemiAnalysis 認為,這對希望服務更廣 interactivity 範圍使用者的供應商而言是一項問題。
Frame:ATTRIBUTE· Mode:ASSERTED· Mapping:COMPLETE
開啟逐條審核

Source: SemiAnalysis InferenceX ↗
GPT-OSS 120B Single Node
MI300X, MI325X, H200, and H100 group in the lower-left of the throughput vs interactivity plot, indicating broadly similar tradeoffs, with Nvidia generally holding a modest lead. The next step up is MI355X, which delivers roughly more than 2x higher token throughput per GPU at a given interactivity level, relative to that first group. Within MI355X, ATOM shifts the curve toward higher throughput at low interactivity, suggesting it prioritizes peak throughput over per-user responsiveness.
Atomic Claim 492/574 · 2026-02-16_inferencex-v2-nvidia-blackwell-vs::IX2-0515
Atomic Claim 493/574 · 2026-02-16_inferencex-v2-nvidia-blackwell-vs::IX2-0516
Atomic Claim 494/574 · 2026-02-16_inferencex-v2-nvidia-blackwell-vs::IX2-0517
Above that tier sits NVIDIA’s B200 and GB200, which outperform MI355X across the frontier. While B200 and GB200 share the same Blackwell compute die, GB200 achieves a higher throughput–interactivity curve because the platform and serving stack reduce non-compute bottlenecks at scale (interconnect/topology, CPU-GPU coupling, and runtime scheduling), translating into effective scale-out and less overhead per token.
Atomic Claim 495/574 · 2026-02-16_inferencex-v2-nvidia-blackwell-vs::IX2-0518
Atomic Claim 496/574 · 2026-02-16_inferencex-v2-nvidia-blackwell-vs::IX2-0519
Atomic Claim 497/574 · 2026-02-16_inferencex-v2-nvidia-blackwell-vs::IX2-0520

Source: SemiAnalysis InferenceX ↗
If we add cost into the equation, MI355x becomes more competitive: beating B200 at high throughputs. However, GB200 still takes the cake for being the cheapest choice.
Atomic Claim 498/574 · 2026-02-16_inferencex-v2-nvidia-blackwell-vs::IX2-0521
Atomic Claim 499/574 · 2026-02-16_inferencex-v2-nvidia-blackwell-vs::IX2-0522

Source: SemiAnalysis InferenceX ↗
Turning again to the comparison between B200 and GB200 NVL72, it is obvious the impact NVL72 has. We discussed the impact of the GB200 NVL72’s larger 72 GPU scale-up world size vs the B200’s 8 GPU scale-up world size earlier in this article. The output token throughput per GPU more than doubles in the ~100 tok/s/user interactivity range, showing the impact of the NVL72’s larger scale up domain.
Atomic Claim 500/574 · 2026-02-16_inferencex-v2-nvidia-blackwell-vs::IX2-0523
Claim: 再次比較 B200 與 GB200 NVL72 時,可以明顯看出 NVL72 架構帶來的影響。
Frame:ATTRIBUTE· Mode:ASSERTED· Mapping:COMPLETE
開啟逐條審核
Atomic Claim 501/574 · 2026-02-16_inferencex-v2-nvidia-blackwell-vs::IX2-0524
Claim: GB200 NVL72 的 scale-up world size 為 72 顆 GPU,相較之下 B200 的 scale-up world size 為 8 顆 GPU。
Frame:COMPARISON· Mode:ASSERTED· Mapping:COMPLETE
開啟逐條審核
Atomic Claim 502/574 · 2026-02-16_inferencex-v2-nvidia-blackwell-vs::IX2-0525

Source: SemiAnalysis InferenceX ↗
Core InferenceX Repo Updates
We have made a few core architectural changes to the InferenceX repository to make it easier to understand and reproduce benchmarks. Additionally, we have fully subscribed to AI usage to maximize productivity and increase developer velocity.
Atomic Claim 503/574 · 2026-02-16_inferencex-v2-nvidia-blackwell-vs::IX2-0526
Claim: SemiAnalysis 對 InferenceX repository 做了數項核心架構調整,以提升 benchmark 的可理解性與可重現性。
Frame:ATTRIBUTE· Mode:ASSERTED· Mapping:COMPLETE
開啟逐條審核
Atomic Claim 504/574 · 2026-02-16_inferencex-v2-nvidia-blackwell-vs::IX2-0527
Claim: SemiAnalysis 也全面導入 AI 工具,以最大化生產力並提高 developer velocity。
Frame:CLAIM_ONLY· Mode:ASSERTED· Mapping:CLAIM_ONLY
開啟逐條審核
Core Changes Since InferenceXv1
One of the main changes we have made since v1 is the cadence with which we perform sweeps. Previously, we were jestermaxing and performed a full sweep over each configuration nightly. However, as we added more chips, disaggregated prefill, wide EP, and other features, we realized that running every single night was way too time consuming and wasteful. Moreover, it’s just not necessary – benchmarks only really need to be re-run when recipes change or a new software version is released.
Atomic Claim 505/574 · 2026-02-16_inferencex-v2-nvidia-blackwell-vs::IX2-0528
Claim: InferenceX v1 之後的一項主要改變,是調整執行 sweeps 的 cadence。
Frame:ATTRIBUTE· Mode:ASSERTED· Mapping:PARTIAL
開啟逐條審核
Atomic Claim 506/574 · 2026-02-16_inferencex-v2-nvidia-blackwell-vs::IX2-0529
Claim: 關於 InferenceX:過去會每晚對所有 configuration 執行一次完整 sweep。
Frame:ATTRIBUTE· Mode:ASSERTED· Mapping:PARTIAL
開啟逐條審核
Atomic Claim 507/574 · 2026-02-16_inferencex-v2-nvidia-blackwell-vs::IX2-0530
Claim: 隨著更多晶片、disaggregated prefill 與 wide EP 被加入測試範圍,sweep 的複雜度提高。
Frame:ATTRIBUTE· Mode:ASSERTED· Mapping:COMPLETE
開啟逐條審核
Atomic Claim 508/574 · 2026-02-16_inferencex-v2-nvidia-blackwell-vs::IX2-0531
Claim: 關於 disaggregated prefill:隨著其他功能持續加入,每晚執行所有 sweep 變得過度耗時且浪費資源。
Frame:ATTRIBUTE· Mode:ASSERTED· Mapping:PARTIAL
開啟逐條審核
Atomic Claim 509/574 · 2026-02-16_inferencex-v2-nvidia-blackwell-vs::IX2-0532
Claim: 關於 InferenceX:Benchmark 實際上只需要在 recipe 改變或新的軟體版本發布時重新執行。
Frame:ATTRIBUTE· Mode:ASSERTED· Mapping:COMPLETE
開啟逐條審核
We now trigger sweeps based on additions to a changelog ↗at the root of the repo. When a developer makes a performance-impacting change to a given config, they add an entry to the changelog listing the affected config along with a brief description of the change. All configs are defined in a master configuration YAML file ↗, which serves as the stateful representation of every data point to be swept, including core settings like ISL/OSL, EP, TP, DP, MTP, and so on. When a PR containing a changelog addition is merged, a workflow parses the referenced config keys, pulls the corresponding sweep definitions from the master config, and fans them out as individual GitHub Actions jobs. The jobs collect all data points for the full sweep and upload the results as artifacts.
Atomic Claim 510/574 · 2026-02-16_inferencex-v2-nvidia-blackwell-vs::IX2-0533
Atomic Claim 511/574 · 2026-02-16_inferencex-v2-nvidia-blackwell-vs::IX2-0535
Atomic Claim 512/574 · 2026-02-16_inferencex-v2-nvidia-blackwell-vs::IX2-0534
Claim: 關於 InferenceX:當開發者對特定 config 做出會影響效能的修改時,會在 changelog 新增一筆紀錄,列出受影響的 config 與變更簡述。
Frame:ATTRIBUTE· Mode:ASSERTED· Mapping:PARTIAL
開啟逐條審核
Atomic Claim 513/574 · 2026-02-16_inferencex-v2-nvidia-blackwell-vs::IX2-0536
Atomic Claim 514/574 · 2026-02-16_inferencex-v2-nvidia-blackwell-vs::IX2-0537
Claim: 關於 InferenceX:這些 jobs 會收集完整 sweep 的所有 data points,並把結果上傳為 artifacts。
Frame:ATTRIBUTE· Mode:ASSERTED· Mapping:PARTIAL
開啟逐條審核
Below is a high-level diagram of how InferenceX launches jobs.
Atomic Claim 515/574 · 2026-02-16_inferencex-v2-nvidia-blackwell-vs::IX2-0538
Claim: 文章提供一張 InferenceX 啟動 jobs 流程的高階架構圖。
Frame:ATTRIBUTE· Mode:ASSERTED· Mapping:COMPLETE
開啟逐條審核

Klaud Cold AI Usage
Shortly after the release of InferenceX v1, we realized how much developer throughput was being left on the table by not utilizing AI more in our InferenceX development. So, we rolled our sleeves up and decided to embrace Claude Code and begin absorbing intelligence, one token at a time to the point that we are currently spending at a 3 million dollars’ worth of Claude intelligence, apply here to join the mission. ↗ We started our enlightenment journey when we realized the GitHub Copilot agent was free – at first we couldn’t believe this feature came at no cost! We soon realized that Copilot is terrible and it became apparent why GitHub was giving it away for free. You probably would have had to _pay us _to keep using it.
Atomic Claim 516/574 · 2026-02-16_inferencex-v2-nvidia-blackwell-vs::IX2-0539
Claim: 在 InferenceX v1 發布不久後,團隊發現未更積極使用 AI,使 InferenceX 開發仍有大量 developer throughput 未被利用。
Frame:ATTRIBUTE· Mode:ASSERTED· Mapping:COMPLETE
開啟逐條審核
Atomic Claim 517/574 · 2026-02-16_inferencex-v2-nvidia-blackwell-vs::IX2-0543
Claim: SemiAnalysis 戲稱,若要繼續使用 Copilot,可能反而需要付費給他們。
Frame:CLAIM_ONLY· Mode:HYPOTHETICAL· Mapping:CLAIM_ONLY
開啟逐條審核
Atomic Claim 518/574 · 2026-02-16_inferencex-v2-nvidia-blackwell-vs::IX2-0540
Claim: 團隊因此全面採用 Claude Code,目前相關支出 run rate 約為每天 6,000 美元。
Frame:ATTRIBUTE· Mode:ASSERTED· Mapping:COMPLETE
開啟逐條審核
Atomic Claim 519/574 · 2026-02-16_inferencex-v2-nvidia-blackwell-vs::IX2-0541
Claim: 團隊最初是因發現 GitHub Copilot agent 免費提供,而開始更深入使用 AI coding tools。
Frame:ATTRIBUTE· Mode:ASSERTED· Mapping:PARTIAL
開啟逐條審核
Atomic Claim 520/574 · 2026-02-16_inferencex-v2-nvidia-blackwell-vs::IX2-0542
We had been using Claude Code locally ever since it was released. But recently, we have integrated Claude Code into InferenceX development, using it for the usual tasks such as reviewing PRs, but we also have given it the ability to perform sweeps on clusters. With the workflows we setup, Claude can manually initiate runs, view the results, and iterate. This has enabled us to deploy quick fixes easily on the go via the GitHub app.
Atomic Claim 521/574 · 2026-02-16_inferencex-v2-nvidia-blackwell-vs::IX2-0544
Claim: 團隊自 Claude Code 發布後便一直在本地端使用它。
Frame:ATTRIBUTE· Mode:ASSERTED· Mapping:PARTIAL
開啟逐條審核
Atomic Claim 522/574 · 2026-02-16_inferencex-v2-nvidia-blackwell-vs::IX2-0545
Claim: 近期團隊已把 Claude Code 整合進 InferenceX 開發流程,並用於 PR review 等一般任務。
Frame:RELATION· Mode:ASSERTED· Mapping:COMPLETE
開啟逐條審核
Atomic Claim 523/574 · 2026-02-16_inferencex-v2-nvidia-blackwell-vs::IX2-0546
Claim: 團隊也讓 Claude Code 能夠在 clusters 上執行 sweeps。
Frame:ATTRIBUTE· Mode:ASSERTED· Mapping:PARTIAL
開啟逐條審核
Atomic Claim 524/574 · 2026-02-16_inferencex-v2-nvidia-blackwell-vs::IX2-0547
Atomic Claim 525/574 · 2026-02-16_inferencex-v2-nvidia-blackwell-vs::IX2-0548
Another cool use case is using Claude to find recipes for new vLLM/SGLang images. When a new image is released, recipes sometimes need to be updated to achieve optimal performance (new environment variables, modified engine arguments, etc.) With our Claude Code integration, we simply open an issue and ask Claude to search through all commits in the image changelog to find necessary changes to be added to the recipe. This works quite well, and although it’s not perfect, it often gives a good starting point.
Atomic Claim 526/574 · 2026-02-16_inferencex-v2-nvidia-blackwell-vs::IX2-0549
Atomic Claim 527/574 · 2026-02-16_inferencex-v2-nvidia-blackwell-vs::IX2-0552
Atomic Claim 528/574 · 2026-02-16_inferencex-v2-nvidia-blackwell-vs::IX2-0551
Claim: 透過 Claude Code 整合,團隊只需建立 issue,要求 Claude 搜尋 image changelog 的所有 commits,以找出 recipe 必須加入的變更。
Frame:RELATION· Mode:ASSERTED· Mapping:COMPLETE
開啟逐條審核
GitHub Actions
In the spirit of open source, all runs occur on GitHub Actions, so benchmark results are verifiable, transparent, and reproducible. However, GitHub outages have been a constant obstacle to our goals recently. We have seen more unicorns lately than any other animal ↗! But maybe it’s time for us to touch some grass.
Atomic Claim 529/574 · 2026-02-16_inferencex-v2-nvidia-blackwell-vs::IX2-0553
Atomic Claim 530/574 · 2026-02-16_inferencex-v2-nvidia-blackwell-vs::IX2-0555
Atomic Claim 531/574 · 2026-02-16_inferencex-v2-nvidia-blackwell-vs::IX2-0554
Atomic Claim 532/574 · 2026-02-16_inferencex-v2-nvidia-blackwell-vs::IX2-0556
Claim: 作者以「也許該去 touch some grass」作為對前述情況的自嘲。
Frame:CLAIM_ONLY· Mode:ASSERTED· Mapping:CLAIM_ONLY
開啟逐條審核
Microsoft/GitHub themselves are aware of this and have stopped updating its status page with aggregate uptime numbers and are down to a single 9: 97.36% over the past 90 days. The problem doesn’t seem to go away if you choose to ignore it…
Atomic Claim 533/574 · 2026-02-16_inferencex-v2-nvidia-blackwell-vs::IX2-0558
Atomic Claim 534/574 · 2026-02-16_inferencex-v2-nvidia-blackwell-vs::IX2-0557
Atomic Claim 535/574 · 2026-02-16_inferencex-v2-nvidia-blackwell-vs::IX2-0559
Claim: 作者諷刺指出,選擇忽略問題並不會讓問題消失。
Frame:CLAIM_ONLY· Mode:INFERRED· Mapping:CLAIM_ONLY
開啟逐條審核

Source: Outages project ↗

Source: Outages project ↗
All in all, GitHub Actions is just alright. It provides a painfully average experience for developers. It is certainly not meant for launching thousands of jobs across a fleet of hundreds of GPUs. Nevertheless, we have worked closely with some GitHub Actions engineers since our launch to better meet the needs of InferenceX, and we can confidently say they have been a pleasure to work with. Moreover, one of our direct asks was to implement lazy loading for jobs when clicking on a workflow run and, while it did take them a while, they eventually implemented the feature. ↗
Atomic Claim 536/574 · 2026-02-16_inferencex-v2-nvidia-blackwell-vs::IX2-0560
Atomic Claim 537/574 · 2026-02-16_inferencex-v2-nvidia-blackwell-vs::IX2-0561
Atomic Claim 538/574 · 2026-02-16_inferencex-v2-nvidia-blackwell-vs::IX2-0562
Atomic Claim 539/574 · 2026-02-16_inferencex-v2-nvidia-blackwell-vs::IX2-0563
Claim: 自 InferenceX 發布以來,團隊持續與部分 GitHub Actions 工程師密切合作,以更符合 InferenceX 的需求。
Frame:RELATION· Mode:ASSERTED· Mapping:COMPLETE
開啟逐條審核
Atomic Claim 540/574 · 2026-02-16_inferencex-v2-nvidia-blackwell-vs::IX2-0564
Atomic Claim 541/574 · 2026-02-16_inferencex-v2-nvidia-blackwell-vs::IX2-0566
Future of InferenceX
Since the initial release of InferenceX in early October 2025, we have worked hard to continuously improve InferenceX. After release, we spent some time refactoring the codebase to make it more scalable, such that new models and inference techniques can now be added in a “plug and play” fashion. These changes enabled us to seamlessly integrate PD-disagg benchmarks for H100, H200, B200, B300, GB200, GB300, and MI355X. We also added accuracy evaluations to our default benchmark pipeline to ensure visibility into model performance across all configurations.
Atomic Claim 542/574 · 2026-02-16_inferencex-v2-nvidia-blackwell-vs::IX2-0567
Claim: 自 InferenceX 於 2025 年 10 月初首次發布後,團隊持續改善 InferenceX。
Frame:ATTRIBUTE· Mode:ASSERTED· Mapping:PARTIAL
開啟逐條審核
Atomic Claim 543/574 · 2026-02-16_inferencex-v2-nvidia-blackwell-vs::IX2-0568
Claim: 關於 InferenceX:發布後,團隊重構 codebase 以提升 scalability,使新的 models 與 inference techniques 可以用「plug and play」方式加入。
Frame:ATTRIBUTE· Mode:ASSERTED· Mapping:COMPLETE
開啟逐條審核
Atomic Claim 544/574 · 2026-02-16_inferencex-v2-nvidia-blackwell-vs::IX2-0569
Atomic Claim 545/574 · 2026-02-16_inferencex-v2-nvidia-blackwell-vs::IX2-0570
Claim: 關於 InferenceX:團隊也把 accuracy evaluations 加入預設 benchmark pipeline,以確保所有 configurations 的 model performance 都可被觀察。
Frame:ATTRIBUTE· Mode:ASSERTED· Mapping:PARTIAL
開啟逐條審核
Although we have made many improvements since our release, there is still much work to be done to achieve the north star goal of providing the most real-world inference benchmarks possible. To achieve this goal, we plan to benchmark on real datasets, add an agentic coding performance benchmark, include more SOTA inference optimizations, benchmark more models, and so much more.
Atomic Claim 546/574 · 2026-02-16_inferencex-v2-nvidia-blackwell-vs::IX2-0571
Claim: 關於 InferenceX:儘管發布後已完成許多改進,但距離提供最貼近真實世界 inference benchmark 的核心目標仍有不少工作。
Frame:ATTRIBUTE· Mode:ASSERTED· Mapping:PARTIAL
開啟逐條審核
Atomic Claim 547/574 · 2026-02-16_inferencex-v2-nvidia-blackwell-vs::IX2-0573
Claim: 為達成此目標,團隊計畫使用真實 datasets、加入 agentic coding performance benchmark,並 benchmark 更多 models。
Frame:ATTRIBUTE· Mode:ASSERTED· Mapping:PARTIAL
開啟逐條審核
Atomic Claim 548/574 · 2026-02-16_inferencex-v2-nvidia-blackwell-vs::IX2-0574
Claim: 為達成此目標,團隊計畫使用真實 datasets、加入 agentic coding performance benchmark,並持續加入更多 benchmark 能力與內容。
Frame:ATTRIBUTE· Mode:ASSERTED· Mapping:PARTIAL
開啟逐條審核
Atomic Claim 549/574 · 2026-02-16_inferencex-v2-nvidia-blackwell-vs::IX2-0572
Claim: 為達成此目標,團隊計畫使用真實 datasets 進行 benchmark、加入 agentic coding performance benchmark,並納入更多 SOTA inference optimizations。
Frame:ATTRIBUTE· Mode:ASSERTED· Mapping:PARTIAL
開啟逐條審核
Migration to Multi Turn Real Multi-Turn Chat and Agentic Coding Datasets
Currently, InferenceX uses completely random tokens as input for benchmarking. We then vary the ISL/OSL uniformly subject to the distribution [ISL*0.8, ISL], similarly for OSL. Because of the random data, we disable prefix caching in all our benchmarks, as the expected value of a prefix cache hit rate on completely random data is 0%. Furthermore, all the random data is single-turn, meaning each conversation contains only one prompt and one response. While this provides a good baseline Pareto frontier, it is not a practical benchmark setup that mimics real-world production inference workloads.
Atomic Claim 550/574 · 2026-02-16_inferencex-v2-nvidia-blackwell-vs::IX2-0575
Claim: 目前 InferenceX 使用完全隨機的 tokens 作為 benchmark input。
Frame:ATTRIBUTE· Mode:ASSERTED· Mapping:PARTIAL
開啟逐條審核
Atomic Claim 551/574 · 2026-02-16_inferencex-v2-nvidia-blackwell-vs::IX2-0577
Claim: 由於使用隨機資料,所有 benchmark 都停用 prefix caching,因為完全隨機資料的 prefix cache 預期 hit rate 為 0%。
Frame:ATTRIBUTE· Mode:ASSERTED· Mapping:PARTIAL
開啟逐條審核
In the near term, we will create a basic multi-turn benchmark with a dataset like allenai/WildChat-4.8M ↗, which captures real users’ multi-turn conversations. In addition to enabling prefix caching on all scenarios, we will enable KV cache CPU offloading, as this is what we see being done in production workloads. This will more accurately evaluate the strengths and weaknesses of each chip. For instance, MI355X has 288GB HBM3e versus B200s 192GB. Therefore, we expect MI355X to perform better in a high concurrency multiturn scenarios as more memory can be allocated to the KV cache. On the other hand, in scenarios where the GPU KV cache is stressed and blocks are offloaded to the CPU, we expect the GBs to excel as these chips have 900GB/s bidirectional CPU-GPU bandwidth, compared to 128GB/s / 256GB/s on HGX with PCIe 5.0 and 6.0, respectively. Moreover, currently we see AMD’s software for CPU offloading is poor, which may negatively affect performance in the same scenarios.
Atomic Claim 552/574 · 2026-02-16_inferencex-v2-nvidia-blackwell-vs::IX2-0581
Claim: 除了在所有情境啟用 prefix caching 外,團隊也計畫啟用 KV cache CPU offloading,因為這是 production workloads 中實際採用的做法。
Frame:NARY_RELATION· Mode:ASSERTED· Mapping:COMPLETE
開啟逐條審核
Atomic Claim 553/574 · 2026-02-16_inferencex-v2-nvidia-blackwell-vs::IX2-0583
Atomic Claim 554/574 · 2026-02-16_inferencex-v2-nvidia-blackwell-vs::IX2-0584
Atomic Claim 555/574 · 2026-02-16_inferencex-v2-nvidia-blackwell-vs::IX2-0588
Atomic Claim 556/574 · 2026-02-16_inferencex-v2-nvidia-blackwell-vs::IX2-0587
Atomic Claim 557/574 · 2026-02-16_inferencex-v2-nvidia-blackwell-vs::IX2-0585
Atomic Claim 558/574 · 2026-02-16_inferencex-v2-nvidia-blackwell-vs::IX2-0586
Atomic Claim 559/574 · 2026-02-16_inferencex-v2-nvidia-blackwell-vs::IX2-0589
The point is: real-world multiturn datasets test more SOTA inference engine features and can capture more nuanced and robust performance data across all chips.
Atomic Claim 560/574 · 2026-02-16_inferencex-v2-nvidia-blackwell-vs::IX2-0591
Claim: 真實世界 multi-turn datasets 能測試更多 SOTA inference engine,並針對所有晶片取得更細緻且更具 robustness 的效能資料。
Frame:ATTRIBUTE· Mode:ASSERTED· Mapping:COMPLETE
開啟逐條審核
Atomic Claim 561/574 · 2026-02-16_inferencex-v2-nvidia-blackwell-vs::IX2-0590
Claim: 真實世界 multi-turn datasets 會測試更多 SOTA inference engine 功能。
Frame:ATTRIBUTE· Mode:ASSERTED· Mapping:COMPLETE
開啟逐條審核
With the rise of Claude Code, Codex, and Kimi, it is becoming increasingly important to benchmark performance in agentic coding scenarios. Like above, these scenarios are multi-turn but also include extremely long context conversations as well as tool use. In the next few months, we plan on creating a benchmark suite that will most accurately capture the performance of open models in these agentic coding scenarios across all chips.
Atomic Claim 562/574 · 2026-02-16_inferencex-v2-nvidia-blackwell-vs::IX2-0592
Claim: 隨著 Claude Code 與 Codex 等工具興起,agentic coding workload 的重要性增加。
Frame:ATTRIBUTE· Mode:ASSERTED· Mapping:COMPLETE
開啟逐條審核
Atomic Claim 563/574 · 2026-02-16_inferencex-v2-nvidia-blackwell-vs::IX2-0593
Claim: Kimi 等工具的興起,使 agentic coding 情境的 benchmark 日益重要。
Frame:ATTRIBUTE· Mode:ASSERTED· Mapping:COMPLETE
開啟逐條審核
Atomic Claim 564/574 · 2026-02-16_inferencex-v2-nvidia-blackwell-vs::IX2-0595
Claim: 未來幾個月,團隊計畫建立 benchmark suite,以更準確衡量各種晶片在 agentic coding 情境下執行 open models 的效能。
Frame:ATTRIBUTE· Mode:ASSERTED· Mapping:COMPLETE
開啟逐條審核
Adding TPU, Trainium and More Models
Currently, we continuously benchmark DeepSeek R1 and GPT OSS 120B (previously Llama 3.1 70B as well). To keep up with the newest model architectures, we plan on adding DeepSeek V3.2 (w/ DSA), DeepSeek V4 on Day 0, Kimi K2.5, Qwen3, GLM5, and many more over the course of the next few months. We will also eventually add multi-modal models and be using EPD & CFD (invented by TogetherAI) optimization too.
Atomic Claim 565/574 · 2026-02-16_inferencex-v2-nvidia-blackwell-vs::IX2-0596
Atomic Claim 566/574 · 2026-02-16_inferencex-v2-nvidia-blackwell-vs::IX2-0597
In addition to new models, we are actively working on adding both TPU and Trainium.
Atomic Claim 567/574 · 2026-02-16_inferencex-v2-nvidia-blackwell-vs::IX2-0599
Total Cost of Ownership (NVL72, Blackwell, Blackwell Ultra, MI355, Hopper, MI325, MI300)
Looking at capital costs across comparable generations, Nvidia systems tend to have higher capital cost than AMD systems. This is driven mostly by higher compute tray content which is driven by higher GPU pricing – it is well known from their financials that Nvidia enjoys higher margins on their GPUs than other vendors. As an example, MI300X compute tray content sits at ~170K for H100 SXM, and the gap widens further in later generations. MI355X is at ~264K and B300 to ~$344K. That incremental silicon content flows directly into higher server cost, and ultimately higher all-in cluster capex per server.
Atomic Claim 568/574 · 2026-02-16_inferencex-v2-nvidia-blackwell-vs::IX2-0600
Atomic Claim 569/574 · 2026-02-16_inferencex-v2-nvidia-blackwell-vs::IX2-0601
Claim: 主要原因是 compute tray content 較高,而這又來自較高的 GPU 價格;從財務資料可知,Nvidia 的 GPUs 毛利率高於其他廠商。
Frame:COMPARISON· Mode:ASSERTED· Mapping:COMPLETE
開啟逐條審核
Atomic Claim 570/574 · 2026-02-16_inferencex-v2-nvidia-blackwell-vs::IX2-0602
Claim: 例如,MI300X compute tray content 約為 13.8 萬美元,而 H100 SXM 約為 17 萬美元,且後續世代差距進一步擴大。
Frame:COMPARISON· Mode:ASSERTED· Mapping:COMPLETE
開啟逐條審核
Atomic Claim 571/574 · 2026-02-16_inferencex-v2-nvidia-blackwell-vs::IX2-0603
This similar dynamic carries over into the Blackwell generation, where an increase in GPU content drives rising total Server cost, in turn driving higher total upfront cluster capex per server, resulting in higher capital cost of ownership.
Atomic Claim 572/574 · 2026-02-16_inferencex-v2-nvidia-blackwell-vs::IX2-0605

Source: SemiAnalysis AI TCO Model ↗
Across comparable generations, operating costs per GPU are broadly similar because chip TDP is the dominant driver of TCO for operating costs. This goes up as you move from H100s to GB300s, given chip TDP double, driving up operating costs per hour per GPU
Atomic Claim 573/574 · 2026-02-16_inferencex-v2-nvidia-blackwell-vs::IX2-0606
Atomic Claim 574/574 · 2026-02-16_inferencex-v2-nvidia-blackwell-vs::IX2-0607

Source: SemiAnalysis AI TCO Model ↗