SA Article Coverage Review · 2026-03-24_nvidia-the-inference-kingdom-expands
Coverage Summary
- Source: 開啟原始 SA 文章
- Atomic Claims:
372 - Source blocks:
227 - Blocks with ≥1 Atomic Claim:
105 - Blocks without Atomic Claim:
122 - Unplaced Claims:
0
Coverage Review
請從頭到尾閱讀下方 SA 全文。Atomic Claim 會依 Evidence 在原文出現的位置 inline 插入。
若該英文段落已有繁中翻譯 cache,翻譯只會作為淡色閱讀輔助顯示;不會進入 source、Claim provenance 或 Graph。
沒有 Claim callout 的段落不一定有問題;若內容重要且應形成知識,請記到 Missing Claim Notes。
Missing Claim Notes
- 若看到重要但沒有 Atomic Claim 的段落,請在這裡記錄:
- Section:
- Evidence:
- 為什麼重要/應該抽成什麼 Claim:
SA Full Text + Translation + Atomic Claims
Nvidia – The Inference Kingdom Expands

Source: Nvidia
At GTC 2026, Nvidia delivered an event packed full of ground breaking announcements. Nvidia’s pace of innovation is not showing any signs of slowing, as they introduced three entirely new systems this year: Groq LPX, Vera ETL256, and STX. Also announced were updates to Nvidia’s Kyber rack architecture system, CPO making its debut for scale-up networking with the unveiling of the Rubin Ultra NVL576 and Feynman NVL1152 multi-rack systems. Early hints on Feynman’s architecture was also a key topic. A Jensen callout for InferenceX during the keynote was a highlight. ↗
Atomic Claim 1/372 · 2026-03-24_nvidia-the-inference-kingdom-expands::NIEK2-0001
Atomic Claim 2/372 · 2026-03-24_nvidia-the-inference-kingdom-expands::NIEK2-0002
Atomic Claim 3/372 · 2026-03-24_nvidia-the-inference-kingdom-expands::NIEK2-0004
Atomic Claim 4/372 · 2026-03-24_nvidia-the-inference-kingdom-expands::NIEK2-0003
Claim: Nvidia 推出了 Vera ETL256。
Frame:RELATION· Mode:ASSERTED· Mapping:COMPLETE
開啟逐條審核
Atomic Claim 5/372 · 2026-03-24_nvidia-the-inference-kingdom-expands::NIEK2-0005
Claim: Nvidia 公布了 Kyber rack 架構系統的更新。
Frame:RELATION· Mode:ASSERTED· Mapping:COMPLETE
開啟逐條審核
Atomic Claim 6/372 · 2026-03-24_nvidia-the-inference-kingdom-expands::NIEK2-0006
Claim: CPO 首度被用於 Rubin Ultra NVL576 多機櫃系統的 scale-up networking。
Frame:NARY_RELATION· Mode:ASSERTED· Mapping:COMPLETE
開啟逐條審核
Atomic Claim 7/372 · 2026-03-24_nvidia-the-inference-kingdom-expands::NIEK2-0007
Claim: CPO 首度被用於 Feynman NVL1152 多機櫃系統的 scale-up networking。
Frame:NARY_RELATION· Mode:ASSERTED· Mapping:PARTIAL
開啟逐條審核
Atomic Claim 8/372 · 2026-03-24_nvidia-the-inference-kingdom-expands::NIEK2-0008
Atomic Claim 9/372 · 2026-03-24_nvidia-the-inference-kingdom-expands::NIEK2-0009
Claim: Jensen 在主題演講中特別提到 InferenceX,成為一項亮點。
Frame:ATTRIBUTE· Mode:ASSERTED· Mapping:PARTIAL
開啟逐條審核
This is our GTC 2026 recap, and we will address many of the key questions that have been left unanswered by Nvidia. Specifically, we will go through the LPX rack and LP30 chip and explain how attention and feed forward network disaggregation (AFD) works; more details on the various rack architectures behind NVL144, NVL576, and NVL1152 and clarify just how much optics will be inserted as well as the rationale behind the dense Vera ETL256. The next generation Kyber rack had some big updates and some hidden details.
Atomic Claim 10/372 · 2026-03-24_nvidia-the-inference-kingdom-expands::NIEK2-0010
Claim: 下一代 Kyber rack 有多項重大更新與一些尚未公開的細節。
Frame:ATTRIBUTE· Mode:ASSERTED· Mapping:COMPLETE
開啟逐條審核
Groq
First up is the Groq LPU. One of the most significant recent events in AI infrastructure was Nvidia’s “acquisition” of Groq. Strictly speaking, Nvidia paid Groq $20B to license their IP and hire most the team. This functions almost as an acquisition, though its structure technically falls short of it being legally considered as one, thereby simplifying or obviating the need for regulatory approvals. Given Nvidia’s market share, if this transaction were structured as a full acquisition and were put to anti-trust review, such a transaction would likely not go through. The other benefit is that it avoids a drawn-out transaction closing process. Nvidia got instant access to Groq’s IP and people. This is why, less than four months after the deal was announced, Nvidia already has a system concept that is being integrated into the Vera Rubin inference stack.
Atomic Claim 11/372 · 2026-03-24_nvidia-the-inference-kingdom-expands::NIEK2-0011
Atomic Claim 12/372 · 2026-03-24_nvidia-the-inference-kingdom-expands::NIEK2-0012
Atomic Claim 13/372 · 2026-03-24_nvidia-the-inference-kingdom-expands::NIEK2-0013
Atomic Claim 14/372 · 2026-03-24_nvidia-the-inference-kingdom-expands::NIEK2-0014
Atomic Claim 15/372 · 2026-03-24_nvidia-the-inference-kingdom-expands::NIEK2-0015
Atomic Claim 16/372 · 2026-03-24_nvidia-the-inference-kingdom-expands::NIEK2-0016
Atomic Claim 17/372 · 2026-03-24_nvidia-the-inference-kingdom-expands::NIEK2-0017
Claim: 因此在交易宣布不到四個月後,Nvidia 已提出一個正整合進 Vera Rubin 推論堆疊的系統概念。
Frame:RELATION· Mode:ASSERTED· Mapping:PARTIAL
開啟逐條審核
Let’s now go through a refresher on the LPU architecture to see how Groq’s LPU complements Nvidia’s GPU. For more details see our original Groq piece. ↗ The premise from that piece remains unchanged: the standalone Groq LPU system is not economical for serving tokens at scale, but it can serve tokens very quickly which can demand a large market premium. This is the premise behind how LPU fits into a disaggregated decode system.
Atomic Claim 18/372 · 2026-03-24_nvidia-the-inference-kingdom-expands::NIEK2-0018
Atomic Claim 19/372 · 2026-03-24_nvidia-the-inference-kingdom-expands::NIEK2-0019
Atomic Claim 20/372 · 2026-03-24_nvidia-the-inference-kingdom-expands::NIEK2-0020
Claim: 這正是 LPU 適合導入 disaggregated decode 系統的核心邏輯。
Frame:RELATION· Mode:ASSERTED· Mapping:PARTIAL
開啟逐條審核
LPU chip
Groq’s first and only publicly announced LPU architecture was detailed in their ISCA 2020 paper. Unlike typical hardware architectures connecting many general-purpose cores, Groq re-organized the architecture into groups of single-purpose units connecting to other groups of different purposes, and they named the groups “slices.” Between functional units are streaming registers, scratchpad SRAM for functional units to pass data to each other. Groq opted for single-level scratchpad SRAM instead of multi-level memory hierarchy to make the hardware execution deterministic.
Atomic Claim 21/372 · 2026-03-24_nvidia-the-inference-kingdom-expands::NIEK2-0021
Atomic Claim 22/372 · 2026-03-24_nvidia-the-inference-kingdom-expands::NIEK2-0022
Atomic Claim 23/372 · 2026-03-24_nvidia-the-inference-kingdom-expands::NIEK2-0023
Atomic Claim 24/372 · 2026-03-24_nvidia-the-inference-kingdom-expands::NIEK2-0024
Concretely, LPU architecture has VXM slices for vector operations, MEM slices for loading/storing data, SXM slices for tensor shape manipulation, and MXM slices for performing matrix multiplication. Spatially, the slices are laid out horizontally, allowing the data to stream horizontally. Within a slice, instructions are pumped vertically across units. Conceptually, LPU resembles a systolic array that pumps instructions vertically and data horizontally.
Atomic Claim 25/372 · 2026-03-24_nvidia-the-inference-kingdom-expands::NIEK2-0025
Atomic Claim 26/372 · 2026-03-24_nvidia-the-inference-kingdom-expands::NIEK2-0027
Atomic Claim 27/372 · 2026-03-24_nvidia-the-inference-kingdom-expands::NIEK2-0026
Atomic Claim 28/372 · 2026-03-24_nvidia-the-inference-kingdom-expands::NIEK2-0028
Atomic Claim 29/372 · 2026-03-24_nvidia-the-inference-kingdom-expands::NIEK2-0029
Atomic Claim 30/372 · 2026-03-24_nvidia-the-inference-kingdom-expands::NIEK2-0030
Atomic Claim 31/372 · 2026-03-24_nvidia-the-inference-kingdom-expands::NIEK2-0031
Claim: 概念上,LPU 類似 systolic array:指令垂直 pumps,資料則水平流動。
Frame:ATTRIBUTE· Mode:ASSERTED· Mapping:PARTIAL
開啟逐條審核

Source: Groq, SemiAnalysis
The data flow and instruction flow design requires fine-grained pipelining to achieve high performance. Since LPU architecture makes computation deterministic, the compiler can aggressively schedule and overlap instructions to hide latency. The LPU’s use of high bandwidth SRAM and aggressive pipelining are the two main factors that enable LPU’s low latency.
Atomic Claim 32/372 · 2026-03-24_nvidia-the-inference-kingdom-expands::NIEK2-0032
Atomic Claim 33/372 · 2026-03-24_nvidia-the-inference-kingdom-expands::NIEK2-0033
Atomic Claim 34/372 · 2026-03-24_nvidia-the-inference-kingdom-expands::NIEK2-0034
LPU gen 1 was designed on a legacy Global Foundries 14nm process, with Marvell responsible for the chip’s physical design. This was a much more mature node compared to peers when it taped out in 2020, with the incumbent AI chip platforms mostly on TSMC’s N7 platform. This made sense for an early product focused on proving out Groq’s architecture and bringing its inference-centric design to market. The 14nm node was mature, relatively well understood, and suitable for an initial chip where architectural differentiation mattered more than pushing its silicon to the leading edge.
Atomic Claim 35/372 · 2026-03-24_nvidia-the-inference-kingdom-expands::NIEK2-0035
Claim: 第一代 LPU 採用較舊的 Global Foundries 14nm 製程設計。
Frame:NARY_RELATION· Mode:ASSERTED· Mapping:COMPLETE
開啟逐條審核
Atomic Claim 36/372 · 2026-03-24_nvidia-the-inference-kingdom-expands::NIEK2-0036
Claim: Marvell 負責該晶片的 physical design。
Frame:RELATION· Mode:ASSERTED· Mapping:COMPLETE
開啟逐條審核
Atomic Claim 37/372 · 2026-03-24_nvidia-the-inference-kingdom-expands::NIEK2-0037
Atomic Claim 38/372 · 2026-03-24_nvidia-the-inference-kingdom-expands::NIEK2-0038
Atomic Claim 39/372 · 2026-03-24_nvidia-the-inference-kingdom-expands::NIEK2-0039
One of the selling points is that the chip can be manufactured and packaged entirely in the United States compared to their competitors being heavily reliant on the Asia semiconductor supply chain: logic and packaging in Taiwan, with HBM from Korea.
Atomic Claim 40/372 · 2026-03-24_nvidia-the-inference-kingdom-expands::NIEK2-0040
Atomic Claim 41/372 · 2026-03-24_nvidia-the-inference-kingdom-expands::NIEK2-0041
Atomic Claim 42/372 · 2026-03-24_nvidia-the-inference-kingdom-expands::NIEK2-0042
Atomic Claim 43/372 · 2026-03-24_nvidia-the-inference-kingdom-expands::NIEK2-0043
Since then, Groq’s roadmap has stalled due to execution, with no LPU 2 having been shipped. This leaves the Groq LPU looking even more dated against competing roadmaps. What was once a meaningful but still manageable node disadvantage versus 7nm-era peers has widened into a far sharper gap, with all leading accelerator platforms now moving onto 3nm-class processes in 2026.
Atomic Claim 44/372 · 2026-03-24_nvidia-the-inference-kingdom-expands::NIEK2-0044
Atomic Claim 45/372 · 2026-03-24_nvidia-the-inference-kingdom-expands::NIEK2-0045
Atomic Claim 46/372 · 2026-03-24_nvidia-the-inference-kingdom-expands::NIEK2-0046
Atomic Claim 47/372 · 2026-03-24_nvidia-the-inference-kingdom-expands::NIEK2-0047
Atomic Claim 48/372 · 2026-03-24_nvidia-the-inference-kingdom-expands::NIEK2-0048
The follow on Groq LPU 2 was designed for Samsung Foundry’s SF4X node, specifically at Samsung’s Austin fab, allowing them to extend the pitch that Groq is fabricated domestically in the USA. Samsung would also provide support for the back-end design. The choice of Samsung was driven by favorable terms / investment, with Samsung Foundry struggling to find customers for its advanced nodes and missing out on an AI logic customer. Unsurprisingly, Samsung was a key investor in Groq’s subsequent Series D in August 2024, and most recently in September 2025 before the Nvidia “acquisition.”
Atomic Claim 49/372 · 2026-03-24_nvidia-the-inference-kingdom-expands::NIEK2-0049
Claim: Groq LPU 2 原設計採用 Samsung Foundry 的 SF4X 製程。
Frame:NARY_RELATION· Mode:ASSERTED· Mapping:COMPLETE
開啟逐條審核
Atomic Claim 50/372 · 2026-03-24_nvidia-the-inference-kingdom-expands::NIEK2-0050
Atomic Claim 51/372 · 2026-03-24_nvidia-the-inference-kingdom-expands::NIEK2-0051
Atomic Claim 52/372 · 2026-03-24_nvidia-the-inference-kingdom-expands::NIEK2-0052
Atomic Claim 53/372 · 2026-03-24_nvidia-the-inference-kingdom-expands::NIEK2-0053
Atomic Claim 54/372 · 2026-03-24_nvidia-the-inference-kingdom-expands::NIEK2-0054
However, the Groq LPU 2 was never productized because of design issues. The C2C SerDes on the chip couldn’t hit the advertised 112G speed which caused the design to malfunction, as we detailed long ago in the Accelerator model ↗. The third generation Groq LPU is the one that Nvidia will be productizing.
Atomic Claim 55/372 · 2026-03-24_nvidia-the-inference-kingdom-expands::NIEK2-0055
Atomic Claim 56/372 · 2026-03-24_nvidia-the-inference-kingdom-expands::NIEK2-0056
Atomic Claim 57/372 · 2026-03-24_nvidia-the-inference-kingdom-expands::NIEK2-0057
SRAM and Memory Hierarchy
We have written about the role of SRAM in the memory hierarchy, but the quick recap is that SRAM is very fast (low latency and high bandwidth) but this comes at the expense of density and therefore cost.
SRAM machines such as Groq’s LPU therefore enable very fast time to first token and tokens per second per user but at the expense of total throughput, as their limited SRAM capacity quickly gets saturated by weights, with little left over for KVcache that grows as more users are batched. GPUs win for throughput and cost as we have shown. This is why Nvidia has decided to combine these architectures to get the best of both worlds: accelerate parts of decode that are more latency sensitive and are not as memory heavy on a low-latency SRAM-heavy chip like the LPU, while memory hungry attention is performed on GPUs that come with a lot of fast (but not SRAM fast) memory capacity.
Atomic Claim 58/372 · 2026-03-24_nvidia-the-inference-kingdom-expands::NIEK2-0058
Claim: Groq 的 LPU 可實現非常短的 time to first token。
Frame:ATTRIBUTE· Mode:ASSERTED· Mapping:COMPLETE
開啟逐條審核
Atomic Claim 59/372 · 2026-03-24_nvidia-the-inference-kingdom-expands::NIEK2-0059
Atomic Claim 60/372 · 2026-03-24_nvidia-the-inference-kingdom-expands::NIEK2-0060
Atomic Claim 61/372 · 2026-03-24_nvidia-the-inference-kingdom-expands::NIEK2-0061
Atomic Claim 62/372 · 2026-03-24_nvidia-the-inference-kingdom-expands::NIEK2-0062
Atomic Claim 63/372 · 2026-03-24_nvidia-the-inference-kingdom-expands::NIEK2-0063
Atomic Claim 64/372 · 2026-03-24_nvidia-the-inference-kingdom-expands::NIEK2-0064
Atomic Claim 65/372 · 2026-03-24_nvidia-the-inference-kingdom-expands::NIEK2-0065

Source: SemiAnalysis
This brings us to the Groq 3 LPU or LP30, with LPU gen 2 being skipped over. This chip has no Nvidia design involvement. The SerDes issues affecting v2 appear to be fixed. Behind the paywall, we will reveal the SerDes IP vendor which may come as a surprise. Nvidia also announced an LP35 which is a minor refresh of the LP30 which will remain on SF4 and will require a new tapeout. It will incorporate NVFP4 number format but given Nvidia is prioritizing time to market we don’t expect any other drastic design changes.
Atomic Claim 66/372 · 2026-03-24_nvidia-the-inference-kingdom-expands::NIEK2-0066
Atomic Claim 67/372 · 2026-03-24_nvidia-the-inference-kingdom-expands::NIEK2-0067
Atomic Claim 68/372 · 2026-03-24_nvidia-the-inference-kingdom-expands::NIEK2-0068
Atomic Claim 69/372 · 2026-03-24_nvidia-the-inference-kingdom-expands::NIEK2-0069
Atomic Claim 70/372 · 2026-03-24_nvidia-the-inference-kingdom-expands::NIEK2-0070
Atomic Claim 71/372 · 2026-03-24_nvidia-the-inference-kingdom-expands::NIEK2-0071
Atomic Claim 72/372 · 2026-03-24_nvidia-the-inference-kingdom-expands::NIEK2-0072
Atomic Claim 73/372 · 2026-03-24_nvidia-the-inference-kingdom-expands::NIEK2-0073

Source: Nvidia
LPU 3’s near reticle size die layout is very similar to LPU 1. a significant amount of area taken is up by the 500MB of on-chip SRAM, with a very small amount of area dedicated to MatMul cores that offer 1.2 PFLOPs of FP8 compute – a fraction of compute compared to Nvidia GPUs. This compares to LPU 1 with 230MB of SRAM and 750 TFLOPs of INT8, with the performance increase mostly driven by node migration from GF16 to SF4. As a single monolithic die, advanced packaging isn’t required.
Atomic Claim 74/372 · 2026-03-24_nvidia-the-inference-kingdom-expands::NIEK2-0074
Atomic Claim 75/372 · 2026-03-24_nvidia-the-inference-kingdom-expands::NIEK2-0075
Atomic Claim 76/372 · 2026-03-24_nvidia-the-inference-kingdom-expands::NIEK2-0076
Atomic Claim 77/372 · 2026-03-24_nvidia-the-inference-kingdom-expands::NIEK2-0077
Atomic Claim 78/372 · 2026-03-24_nvidia-the-inference-kingdom-expands::NIEK2-0078
Atomic Claim 79/372 · 2026-03-24_nvidia-the-inference-kingdom-expands::NIEK2-0079
One of the benefits of relying on SF4 is that it isn’t constrained like TSMC’s N3, which is putting a cap on accelerator production and is a key reason why the industry remains compute constrained. ↗ This is in addition to not having HBM which is also constrained ↗. This allows Nvidia to ramp production of the LPU without sacrificing or eating into their valuable TSMC allocation or HBM allocations, representing true incremental revenue and capacity that noone else can access.
Atomic Claim 80/372 · 2026-03-24_nvidia-the-inference-kingdom-expands::NIEK2-0081
Atomic Claim 81/372 · 2026-03-24_nvidia-the-inference-kingdom-expands::NIEK2-0080
Atomic Claim 82/372 · 2026-03-24_nvidia-the-inference-kingdom-expands::NIEK2-0082
Atomic Claim 83/372 · 2026-03-24_nvidia-the-inference-kingdom-expands::NIEK2-0083
Since Nvidia has taken over, the next generation LP40 will be fabricated on TSMC N3P and use CoWoS-R, and Nvidia will contribute more of their own IP such as supporting the NVLink protocol rather than Groq’s C2C. This will be the first LPU to be extremely co-designed alongside the Feynman platform. Groq’s original plans for LPU Gen 4 was also with TSMC and Alchip as the back-end design partner. Alchip’s involvement is now redundant with Nvidia able to perform backend design on their own. One of the technical innovations planned is hybrid bonded DRAM to extend on-chip memory with only a slight decrease in latency and bandwidth vs SRAM, but much higher performance compared to DRAM. SK Hynix was tapped as the supplier of the DRAM to be used for the 3D stacking. All of this and more was detailed long ago in the Accelerator model ↗.
Atomic Claim 84/372 · 2026-03-24_nvidia-the-inference-kingdom-expands::NIEK2-0084
Atomic Claim 85/372 · 2026-03-24_nvidia-the-inference-kingdom-expands::NIEK2-0085
Atomic Claim 86/372 · 2026-03-24_nvidia-the-inference-kingdom-expands::NIEK2-0086
Atomic Claim 87/372 · 2026-03-24_nvidia-the-inference-kingdom-expands::NIEK2-0087
Atomic Claim 88/372 · 2026-03-24_nvidia-the-inference-kingdom-expands::NIEK2-0088
Atomic Claim 89/372 · 2026-03-24_nvidia-the-inference-kingdom-expands::NIEK2-0089
Atomic Claim 90/372 · 2026-03-24_nvidia-the-inference-kingdom-expands::NIEK2-0090
Atomic Claim 91/372 · 2026-03-24_nvidia-the-inference-kingdom-expands::NIEK2-0091
Atomic Claim 92/372 · 2026-03-24_nvidia-the-inference-kingdom-expands::NIEK2-0092
Claim: SK Hynix 被選為供應用於 3D stacking 的 DRAM 供應商。
Frame:NARY_RELATION· Mode:ASSERTED· Mapping:PARTIAL
開啟逐條審核
Atomic Claim 93/372 · 2026-03-24_nvidia-the-inference-kingdom-expands::NIEK2-0093
Claim: 上述內容與更多細節,SemiAnalysis 先前已在 Accelerator model 中說明。
Frame:CLAIM_ONLY· Mode:ASSERTED· Mapping:CLAIM_ONLY
開啟逐條審核

Source: Nvidia, SemiAnalysis Accelerator Model ↗
GPU and LPU Integration: Attention FFN Disaggregation (AFD)

Source: Nvidia
Now with an understanding of what LPUs are good for we can understand how they fit into inference setups. NVIDIA introduced LPUs to improve the performance of high interactivity scenarios. In those scenarios, LPUs can leverage their low-latency capabilities to improve the decode phase latencies. One way LPUs can improve decode phase latencies is by applying the Attention FFN Disaggregation (AFD) technique, introduced in MegaScale-Infer ↗ and Step-3 ↗.
Atomic Claim 94/372 · 2026-03-24_nvidia-the-inference-kingdom-expands::NIEK2-0096
Claim: LPU 改善 decode 延遲的一種方式,是採用 Attention FFN Disaggregation(AFD)技術;此技術最早見於 MegaScale-Infer 與 Step-3。
Frame:NARY_RELATION· Mode:ASSERTED· Mapping:PARTIAL
開啟逐條審核
Atomic Claim 95/372 · 2026-03-24_nvidia-the-inference-kingdom-expands::NIEK2-0094
Atomic Claim 96/372 · 2026-03-24_nvidia-the-inference-kingdom-expands::NIEK2-0095
As we explained in our InferenceX article ↗, LLM inference involves two phases: prefill and decode. Prefill processes the full input context: It is compute-intensive, which is suitable for GPUs. On the other hand, decode predicts new tokens and is memory-bounded. Decode is latency-sensitive because the model predicts new tokens one by one, and LPU’s high SRAM bandwidth and low-latency capabilities can help accelerate this iterative process.
Atomic Claim 97/372 · 2026-03-24_nvidia-the-inference-kingdom-expands::NIEK2-0097
Claim: 如 InferenceX 文章所述,LLM 推論分為 prefill 與 decode 兩個階段。
Frame:NARY_RELATION· Mode:ASSERTED· Mapping:PARTIAL
開啟逐條審核
Atomic Claim 98/372 · 2026-03-24_nvidia-the-inference-kingdom-expands::NIEK2-0098
Atomic Claim 99/372 · 2026-03-24_nvidia-the-inference-kingdom-expands::NIEK2-0099
Atomic Claim 100/372 · 2026-03-24_nvidia-the-inference-kingdom-expands::NIEK2-0100
Atomic Claim 101/372 · 2026-03-24_nvidia-the-inference-kingdom-expands::NIEK2-0101

Source: SemiAnalysis
Attention and FFN are subsets of operations in a model. In a model forward pass, attention’s output feeds into a token router, and the token router assigns each token to k experts, where each expert is an FFN. Attention and FFN have very different performance properties. During decode phase, the GPU utilization of attention barely improves when scaling batch size due to being bounded by loading KV cache. In contrast, the GPU utilization of FFN scales with batch size comparatively better.
Atomic Claim 102/372 · 2026-03-24_nvidia-the-inference-kingdom-expands::NIEK2-0102
Atomic Claim 103/372 · 2026-03-24_nvidia-the-inference-kingdom-expands::NIEK2-0103
Atomic Claim 104/372 · 2026-03-24_nvidia-the-inference-kingdom-expands::NIEK2-0104
Atomic Claim 105/372 · 2026-03-24_nvidia-the-inference-kingdom-expands::NIEK2-0105
Atomic Claim 106/372 · 2026-03-24_nvidia-the-inference-kingdom-expands::NIEK2-0106
Claim: 在 decode 階段,attention 因受到 KV cache 載入限制,即使提高 batch size,GPU 利用率也幾乎不會改善。
Frame:ATTRIBUTE· Mode:ASSERTED· Mapping:COMPLETE
開啟逐條審核
Atomic Claim 107/372 · 2026-03-24_nvidia-the-inference-kingdom-expands::NIEK2-0107
Claim: 相較之下,FFN 的 GPU 利用率會隨 batch size 增加而有較明顯提升。
Frame:COMPARISON· Mode:ASSERTED· Mapping:COMPLETE
開啟逐條審核
This is something we have worked with certain hardware vendors and memory companies on with our inference simulator for more than 6 months. ↗
Atomic Claim 108/372 · 2026-03-24_nvidia-the-inference-kingdom-expands::NIEK2-0108
Claim: SemiAnalysis 已與部分硬體供應商與記憶體公司使用其 inference simulator 研究此議題超過六個月。
Frame:CLAIM_ONLY· Mode:ASSERTED· Mapping:CLAIM_ONLY
開啟逐條審核

Source: MegaScale-Infer, SemiAnalysis
As state-of-the-art mixture-of-expert (MoE) models grow increasingly sparse, tokens can choose experts from a larger expert pool. As a result, each expert receives fewer tokens, leading to lower utilization. This motivates attention and FFN disaggregation. If a GPU only performs attention operations, its HBM capacity can be fully allocated to KV cache, increasing the total number of tokens it can process, which then increases the tokens each expert processes on average.
Atomic Claim 109/372 · 2026-03-24_nvidia-the-inference-kingdom-expands::NIEK2-0109
Atomic Claim 110/372 · 2026-03-24_nvidia-the-inference-kingdom-expands::NIEK2-0110
Atomic Claim 111/372 · 2026-03-24_nvidia-the-inference-kingdom-expands::NIEK2-0111
Atomic Claim 112/372 · 2026-03-24_nvidia-the-inference-kingdom-expands::NIEK2-0112

Source: SemiAnalysis
Comparing the two operations, we see attention is stateful due to dynamic KV cache loading patterns, whereas FFN is stateless since the computation only depends on the token inputs. Thus, we disaggregate the computation of attention and FFN. We map attention computations to GPUs, which handle dynamic workloads well. For FFNs, we map them to LPUs, since LPU architecture is inherently deterministic and benefits from static compute workloads.
Atomic Claim 113/372 · 2026-03-24_nvidia-the-inference-kingdom-expands::NIEK2-0113
Atomic Claim 114/372 · 2026-03-24_nvidia-the-inference-kingdom-expands::NIEK2-0114
Atomic Claim 115/372 · 2026-03-24_nvidia-the-inference-kingdom-expands::NIEK2-0115
Atomic Claim 116/372 · 2026-03-24_nvidia-the-inference-kingdom-expands::NIEK2-0116
Atomic Claim 117/372 · 2026-03-24_nvidia-the-inference-kingdom-expands::NIEK2-0117

Source: SemiAnalysis, MegaScale-Infer
With AFD, token routing from GPUs to LPUs can become the bottleneck, especially under strict latency constraints. The token routing flow involves two operations: dispatch and combine. In the dispatch step, we route each token to their top k experts with an All-to-All collective operation. After experts complete their computation, we perform the combine step, where the outputs are sent back to the source location with a reverse All-to-All collective, continuing the next layer’s computation.
Atomic Claim 118/372 · 2026-03-24_nvidia-the-inference-kingdom-expands::NIEK2-0118
Atomic Claim 119/372 · 2026-03-24_nvidia-the-inference-kingdom-expands::NIEK2-0119
Atomic Claim 120/372 · 2026-03-24_nvidia-the-inference-kingdom-expands::NIEK2-0120
Atomic Claim 121/372 · 2026-03-24_nvidia-the-inference-kingdom-expands::NIEK2-0121

Source: SemiAnalysis
To hide the communication latency of dispatch and combine, we employ ping pong pipeline parallelism. In addition to splitting batches into micro-batches and computation pipelining like standard pipeline parallelism, the tokens dispatched to the LPUs are combined back to the source GPUs, so they ping pong between the GPUs and the LPUs.
Atomic Claim 122/372 · 2026-03-24_nvidia-the-inference-kingdom-expands::NIEK2-0122
Claim: 為隱藏 dispatch 與 combine 的通訊延遲,系統採用 ping-pong pipeline parallelism。
Frame:ATTRIBUTE· Mode:ASSERTED· Mapping:COMPLETE
開啟逐條審核
Atomic Claim 123/372 · 2026-03-24_nvidia-the-inference-kingdom-expands::NIEK2-0123
Claim: 除了像標準 pipeline parallelism 一樣把 batch 切成 micro-batches 並進行運算 pipeline 外,送往 LPU 的 token 還會再 combine 回來源 GPUs,使資料在 GPUs 與 LPU 之間來回傳遞。
Frame:ATTRIBUTE· Mode:ASSERTED· Mapping:PARTIAL
開啟逐條審核

Source: MegaScale-Infer

Source: SemiAnalysis

Source: SemiAnalysis
Speculative Decoding
A different way LPUs could improve decode phase latencies is by accelerating a speculative decoding setup, where we deploy draft models or Multi-Token Prediction (MTP) layers onto LPUs.
Atomic Claim 124/372 · 2026-03-24_nvidia-the-inference-kingdom-expands::NIEK2-0124
Claim: LPU 改善 decode 延遲的另一種可能方式,是加速 speculative decoding,將 draft model 或 Multi-Token Prediction(MTP)layers 部署到 LPU 上。
Frame:NARY_RELATION· Mode:HYPOTHETICAL· Mapping:PARTIAL
開啟逐條審核
For a decoding step of context N tokens, adding k additional tokens during forward pass (a warm prefill of k new tokens) marginally increases the latency when k << N. Using this property, speculative decoding uses a small draft model or MTP layers to predict k new tokens, saving time since small models have lower latency per decode step. To verify the draft tokens, the main model only needs one warm prefill of k new tokens, at the latency cost of roughly a single decode step. Speculative decoding usually boosts output token per decode step by 1.5 to 2 tokens, depending on the draft model / MTP accuracy. With its low latency capabilities, LPUs can further increase the latency savings and improve throughput.
Atomic Claim 125/372 · 2026-03-24_nvidia-the-inference-kingdom-expands::NIEK2-0125
Atomic Claim 126/372 · 2026-03-24_nvidia-the-inference-kingdom-expands::NIEK2-0126
Claim: 利用這項特性,speculative decoding 可由較小的 draft model 或 MTP layers 預測 k 個新 token;由於小模型每個 decode step 延遲較低,因此能節省時間。
Frame:COMPARISON· Mode:ASSERTED· Mapping:COMPLETE
開啟逐條審核
Atomic Claim 127/372 · 2026-03-24_nvidia-the-inference-kingdom-expands::NIEK2-0127
Atomic Claim 128/372 · 2026-03-24_nvidia-the-inference-kingdom-expands::NIEK2-0128
Claim: Speculative decoding 通常可將每個 decode step 的輸出 token 數提高至 1.5~2 個,實際效果取決於 draft model/MTP 的準確率。
Frame:ATTRIBUTE· Mode:ASSERTED· Mapping:COMPLETE
開啟逐條審核
Atomic Claim 129/372 · 2026-03-24_nvidia-the-inference-kingdom-expands::NIEK2-0129

Source: SemiAnalysis
For LPUs, deploying a draft model or MTP layers is quite different from applying AFD. FFNs are stateless, while draft models and MTP layers require dynamic KV cache loading. Each FFN is around hundreds of megabytes, whereas draft models and MTP layers take up tens of gigabytes. To support this memory usage, LPUs can access up to 256 GB of DDR5 per Fabric Expansion Logic FPGAs on the LPX compute tray.
Atomic Claim 130/372 · 2026-03-24_nvidia-the-inference-kingdom-expands::NIEK2-0130
Atomic Claim 131/372 · 2026-03-24_nvidia-the-inference-kingdom-expands::NIEK2-0131
Atomic Claim 132/372 · 2026-03-24_nvidia-the-inference-kingdom-expands::NIEK2-0132
Atomic Claim 133/372 · 2026-03-24_nvidia-the-inference-kingdom-expands::NIEK2-0133
Atomic Claim 134/372 · 2026-03-24_nvidia-the-inference-kingdom-expands::NIEK2-0134
Atomic Claim 135/372 · 2026-03-24_nvidia-the-inference-kingdom-expands::NIEK2-0135
Claim: 為支援這些記憶體需求,LPX compute tray 上每顆 Fabric Expansion Logic FPGAs 最多可讓 LPU 存取 256GB DDR5。
Frame:NARY_RELATION· Mode:ASSERTED· Mapping:COMPLETE
開啟逐條審核
LPX Rack System
Let’s look at the LPX rack system, which has interesting details. Nvidia has displayed an LPX rack with 32 1U LPU compute trays with 2 Spectrum-X switches. This 32 tray 1U version that Nvidia has shown off at GTC is very close to Groq’s original server design before the acquisition. We believe that this server configuration is not the version that will be shipped in 3Q, with Nvidia implementing changes. Here, we will detail what we know about the actual production version. This was already detailed in the Accelerator model ↗.
Atomic Claim 136/372 · 2026-03-24_nvidia-the-inference-kingdom-expands::NIEK2-0136
Atomic Claim 137/372 · 2026-03-24_nvidia-the-inference-kingdom-expands::NIEK2-0137
Atomic Claim 138/372 · 2026-03-24_nvidia-the-inference-kingdom-expands::NIEK2-0138
Atomic Claim 139/372 · 2026-03-24_nvidia-the-inference-kingdom-expands::NIEK2-0139
Claim: 這些內容先前已在 Accelerator model 中說明。
Frame:CLAIM_ONLY· Mode:ASSERTED· Mapping:CLAIM_ONLY
開啟逐條審核

Source: SemiAnalysis Accelerator Model ↗
LPX Compute Tray
Each LPX compute tray or node has 16 LPUs with 2 Altera FPGAs, 1 Intel Granite Rapids host CPU and 1 BlueField-4 front-end module. As with other Nvidia systems, hyperscalers customers can and will use their own Front-end NIC of choice rather than paying for Nvidia’s BlueField.
Atomic Claim 140/372 · 2026-03-24_nvidia-the-inference-kingdom-expands::NIEK2-0142
Claim: 每個 LPX compute tray/node 配備 1 個 BlueField-4 front-end module。
Frame:NARY_RELATION· Mode:ASSERTED· Mapping:COMPLETE
開啟逐條審核
Atomic Claim 141/372 · 2026-03-24_nvidia-the-inference-kingdom-expands::NIEK2-0141
Claim: 每個 LPX compute tray/node 配備 1 顆 Intel Granite Rapids host CPU。
Frame:NARY_RELATION· Mode:ASSERTED· Mapping:COMPLETE
開啟逐條審核
Atomic Claim 142/372 · 2026-03-24_nvidia-the-inference-kingdom-expands::NIEK2-0140
Claim: 每個 LPX compute tray/node 配備 16 顆 LPU 與 2 顆 Altera FPGAs。
Frame:NARY_RELATION· Mode:ASSERTED· Mapping:COMPLETE
開啟逐條審核
Atomic Claim 143/372 · 2026-03-24_nvidia-the-inference-kingdom-expands::NIEK2-0143
Claim: 和其他 Nvidia 系統相同,hyperscalers 客戶可以自行選擇前端網路方案。
Frame:COMPARISON· Mode:ASSERTED· Mapping:PARTIAL
開啟逐條審核
Atomic Claim 144/372 · 2026-03-24_nvidia-the-inference-kingdom-expands::NIEK2-0144
Claim: hyperscalers 客戶會選用自己的 Front-end NIC,而非支付額外成本採用 Nvidia 的 BlueField。
Frame:NARY_RELATION· Mode:ASSERTED· Mapping:COMPLETE
開啟逐條審核

Source: SemiAnalysis Accelerator Model ↗
The LPU modules are mounted in a belly-to-belly on the PCB, meaning 8 LP30 modules on the top side of the PCB and the other 8 LP30 modules on the bottom. All of the connectivity that comes out of the LPU are via PCB traces and given the dense all-to-all mesh for intra-node connections this requires a very high spec PCB to support the routing. The belly-to-belly mounting is used to reduce PCB trace lengths across the ‘X’ and ‘Y’ dimensions.
Atomic Claim 145/372 · 2026-03-24_nvidia-the-inference-kingdom-expands::NIEK2-0145
Atomic Claim 146/372 · 2026-03-24_nvidia-the-inference-kingdom-expands::NIEK2-0146
Atomic Claim 147/372 · 2026-03-24_nvidia-the-inference-kingdom-expands::NIEK2-0147

Source: SemiAnalysis Networking Model ↗
Something interesting about the system is the important role the FPGAs play. Nvidia refers to the FPGAs as “Fabric Expansion Logic” which serves multiple purposes. First, they act as a NIC which converts the LPU’s C2C protocol into Ethernet to connect to the Spectrum-X based ethernet scale-out fabric. It is this scale-out fabric through which the LPUs connect to GPUs in the decode system.
Atomic Claim 148/372 · 2026-03-24_nvidia-the-inference-kingdom-expands::NIEK2-0148
Atomic Claim 149/372 · 2026-03-24_nvidia-the-inference-kingdom-expands::NIEK2-0149
Atomic Claim 150/372 · 2026-03-24_nvidia-the-inference-kingdom-expands::NIEK2-0150
Atomic Claim 151/372 · 2026-03-24_nvidia-the-inference-kingdom-expands::NIEK2-0151
Second, the LPUs also traverse through the FPGAs to reach the host CPU, with the FPGAs converting C2C to PCIe to the CPU.
Atomic Claim 152/372 · 2026-03-24_nvidia-the-inference-kingdom-expands::NIEK2-0152
Third, the FPGAs are connected to the backplane to talk to other FPGAs in the node, we believe this is to help manage control flow and timing of all the LPUs. The FPGAs also bring extra system DRAM of up to 256GB each. This pool of memory can be used for KVCache if the user wants the entire decode process served by the LPX.
Atomic Claim 153/372 · 2026-03-24_nvidia-the-inference-kingdom-expands::NIEK2-0153
Atomic Claim 154/372 · 2026-03-24_nvidia-the-inference-kingdom-expands::NIEK2-0154
Atomic Claim 155/372 · 2026-03-24_nvidia-the-inference-kingdom-expands::NIEK2-0155
On the front panel there are 8 x OSFP cages for cross-rack C2C, while there will be 2 cages (likely QSFP-DD) that goes to the Spectrum-switches that is used to connect the LPUs and the GPUs for the disaggregated decode system. We will share more about this when we describe the network.
Atomic Claim 156/372 · 2026-03-24_nvidia-the-inference-kingdom-expands::NIEK2-0156
Atomic Claim 157/372 · 2026-03-24_nvidia-the-inference-kingdom-expands::NIEK2-0157
Claim: 另有 2 個 cages(很可能是 QSFP-DD)連到 Spectrum switches,用來在 disaggregated decode 系統中連接 LPU 與 GPUs。
Frame:NARY_RELATION· Mode:EXPECTED· Mapping:PARTIAL
開啟逐條審核
LPU Network
The LPU network can be divided into the scale-up ‘C2C’ network and scale-out network which interacts with the Nvidia GPUs through Spectrum-X. First let’s discuss the scale-up network which can be divided into 3 portions: intra-node, inter-node/intra-rack, inter-rack. For C2C within the rack Nvidia announced a total of 640TB/s of scale up bandwidth per rack which comes from 256 LPUs x 90 lanes x 112Gbps/8 x 2 directions = 645TB/s. Note that Nvidia uses the total 112G line rate rather than 100G of effective data rate.
Atomic Claim 158/372 · 2026-03-24_nvidia-the-inference-kingdom-expands::NIEK2-0158
Atomic Claim 159/372 · 2026-03-24_nvidia-the-inference-kingdom-expands::NIEK2-0159
Claim: scale-up network 又可分成 intra-node、inter-node/intra-rack、inter-rack 三個部分。
Frame:ATTRIBUTE· Mode:ASSERTED· Mapping:COMPLETE
開啟逐條審核
Atomic Claim 160/372 · 2026-03-24_nvidia-the-inference-kingdom-expands::NIEK2-0160
Atomic Claim 161/372 · 2026-03-24_nvidia-the-inference-kingdom-expands::NIEK2-0161
Intra-Tray Topology

Source: SemiAnalysis Networking Model ↗
Within each tray or node, all 16 LPUs are connected to each other in an all-to-all mesh. Each LPU module connects to the 15 other LPUs within the node with 4x100G of C2C bandwidth. Note that this ‘C2C’ is not related to NVLink, but Groq’s own scaleup fabric. These connections are all via PCB trace, which necessitates an extremely high spec PCB to support this routing density. This is why the belly-to-belly layout is used: it reduces the ‘X’ and ‘Y’ distance between all the LPUs and instead have routing go in the ‘Z’ dimension.
Atomic Claim 162/372 · 2026-03-24_nvidia-the-inference-kingdom-expands::NIEK2-0162
Atomic Claim 163/372 · 2026-03-24_nvidia-the-inference-kingdom-expands::NIEK2-0163
Atomic Claim 164/372 · 2026-03-24_nvidia-the-inference-kingdom-expands::NIEK2-0164
Atomic Claim 165/372 · 2026-03-24_nvidia-the-inference-kingdom-expands::NIEK2-0165
Atomic Claim 166/372 · 2026-03-24_nvidia-the-inference-kingdom-expands::NIEK2-0166
The LPU also has 1x100G going to one FPGA, with each FPGA interfacing with 8 LPUs. The 2 FPGAs each have 8x PCIe Gen 5 going to the CPUs. The LPU needs to traverse through the FPGA to interface with the CPU as LPUs don’t have PCIe PHYs to interface directly.
Atomic Claim 167/372 · 2026-03-24_nvidia-the-inference-kingdom-expands::NIEK2-0167
Atomic Claim 168/372 · 2026-03-24_nvidia-the-inference-kingdom-expands::NIEK2-0168
Atomic Claim 169/372 · 2026-03-24_nvidia-the-inference-kingdom-expands::NIEK2-0169
Inter-node/Intra-rack

Source: SemiAnalysis Networking Model ↗
Each LPU connects to one LPU from each of the 15 other nodes in the server. Each of these inter-node links is 2x100G so there are 15x2x100G inter-node links coming out of each LPU. These inter-node links are via a copper cable backplane. In addition, each FPGA also connects to an FPGA in every other node at either 25G or 50G per link for 15x25G/50G. This also goes through the backplane. This means that each node has 16 x 15 x 2 lanes for inter-node C2C and 2 x 15 lanes for inter-node FPGA which is a total of 510 lanes or 1020 differential pairs (for Rx and Tx). Therefore, the backplane is 16 x 1020/2 = 8,160 differential pairs – we divide by 2 as each device Tx channel is a corresponding device’s Rx channel.
Atomic Claim 170/372 · 2026-03-24_nvidia-the-inference-kingdom-expands::NIEK2-0170
Atomic Claim 171/372 · 2026-03-24_nvidia-the-inference-kingdom-expands::NIEK2-0171
Atomic Claim 172/372 · 2026-03-24_nvidia-the-inference-kingdom-expands::NIEK2-0172
Claim: 這些 inter-node links 透過 copper cable backplane 傳輸。
Frame:ATTRIBUTE· Mode:ASSERTED· Mapping:PARTIAL
開啟逐條審核
Atomic Claim 173/372 · 2026-03-24_nvidia-the-inference-kingdom-expands::NIEK2-0173
Atomic Claim 174/372 · 2026-03-24_nvidia-the-inference-kingdom-expands::NIEK2-0174
Atomic Claim 175/372 · 2026-03-24_nvidia-the-inference-kingdom-expands::NIEK2-0175
Atomic Claim 176/372 · 2026-03-24_nvidia-the-inference-kingdom-expands::NIEK2-0176
Inter-rack

Source: SemiAnalysis Networking Model ↗
Lastly, there is the inter-rack C2C. Each LPU has 4x100G lanes that go to the OSFP cages to connect LPUs across 4 racks. There are various configurations that can be used for this inter-rack scale up. One option is 4x100G from each LPU going to one OSFP cage, each OSFP escaping 800G of C2C from 2 LPUs. However, for greater fan out the preferred configuration seems to be each 100G lane from the LPU going to 4 individual cages, with each cage escaping 800G of C2C from 8 LPUs. In terms of how the racks are networked together it appears to be a daisy chain configuration, with each Node0 connected to 2 other Node 0. This can all be achieved within the reach of 100G AECs, though optics can be used if necessary.
Atomic Claim 177/372 · 2026-03-24_nvidia-the-inference-kingdom-expands::NIEK2-0177
Atomic Claim 178/372 · 2026-03-24_nvidia-the-inference-kingdom-expands::NIEK2-0178
Atomic Claim 179/372 · 2026-03-24_nvidia-the-inference-kingdom-expands::NIEK2-0179
Atomic Claim 180/372 · 2026-03-24_nvidia-the-inference-kingdom-expands::NIEK2-0180
Atomic Claim 181/372 · 2026-03-24_nvidia-the-inference-kingdom-expands::NIEK2-0181
Atomic Claim 182/372 · 2026-03-24_nvidia-the-inference-kingdom-expands::NIEK2-0182
Atomic Claim 183/372 · 2026-03-24_nvidia-the-inference-kingdom-expands::NIEK2-0183
Atomic Claim 184/372 · 2026-03-24_nvidia-the-inference-kingdom-expands::NIEK2-0184
Nvidia’s CPO Roadmap
NVIDIA revealed its CPO Roadmap at the GTC Keynote 2026, with Jensen following up with additional commentary in the Financial Analyst Q+A meeting held the following day. Though many had their hopes up for CPO to be used for scale-up within the rack for Rubin Ultra Kyber, Nvidia’s focus was instead on using CPO to enable larger world size compute systems.
Atomic Claim 185/372 · 2026-03-24_nvidia-the-inference-kingdom-expands::NIEK2-0185
Atomic Claim 186/372 · 2026-03-24_nvidia-the-inference-kingdom-expands::NIEK2-0186

Source: SemiAnalysis AI Networking Model ↗, Nvidia
In the Rubin Generation, Nvidia will offer the Rubin GPU in an Oberon NVL72 form factor with an all-copper scale-up network. For Rubin Ultra, as we expected, there will only be a copper scale-up option for Rubin Ultra in the Oberon and Kyber Rack form factor. Rubin Ultra will also be offered in a larger world size system that connects 8 Oberon Racks of 72 Rubin Ultra GPUs to form what will be referred to as NVL576. CPO scale-up will be used to build the larger world size, connecting between the racks in a two-tier all to all network, though scale-up inside the racks will remain copper-based.
Atomic Claim 187/372 · 2026-03-24_nvidia-the-inference-kingdom-expands::NIEK2-0187
Claim: 在 Rubin 世代,Nvidia 將以 Oberon NVL72 form factor 提供 Rubin GPU,且 scale-up network 全部採用 copper。
Frame:NARY_RELATION· Mode:ASSERTED· Mapping:PARTIAL
開啟逐條審核
Atomic Claim 188/372 · 2026-03-24_nvidia-the-inference-kingdom-expands::NIEK2-0188
Claim: 對 Rubin Ultra 而言,在 Oberon 與 Kyber Rack form factor 中,Rubin Ultra 的 scale-up 都只會提供 copper 方案。
Frame:NARY_RELATION· Mode:ASSERTED· Mapping:PARTIAL
開啟逐條審核
Atomic Claim 189/372 · 2026-03-24_nvidia-the-inference-kingdom-expands::NIEK2-0189
Claim: Rubin Ultra 還會提供更大的 world-size 系統,將 8 個 Oberon racks、每櫃 72 顆 Rubin Ultra GPUs 串接成 NVL576。
Frame:NARY_RELATION· Mode:ASSERTED· Mapping:PARTIAL
開啟逐條審核
Atomic Claim 190/372 · 2026-03-24_nvidia-the-inference-kingdom-expands::NIEK2-0190
When we reach the Feynman Generation, CPO usage will expand via another large world size rack, the NVL1152 which is formed by combining 8 Kyber racks. While the Nvidia Technical Blog ↗ that outlines the rack configuration roadmap states that “NVIDIA Kyber will scale up into a massive all-to-all NVL1152 supercomputer using similar direct optical interconnects for rack-to-rack scale-up”, Jensen Huang in a Financial Analyst Q+A session did say that NVL1152 in Feynman would be “all CPO”. There is some disagreement on whether copper will still be used for scale-up within the rack or whether CPO will replace copper.
Atomic Claim 191/372 · 2026-03-24_nvidia-the-inference-kingdom-expands::NIEK2-0191
Claim: 進入 Feynman Generation 後,CPO 應用將進一步擴大至 NVL1152,由 8 個 Kyber racks 組成。
Frame:NARY_RELATION· Mode:ASSERTED· Mapping:COMPLETE
開啟逐條審核
Atomic Claim 192/372 · 2026-03-24_nvidia-the-inference-kingdom-expands::NIEK2-0192
Atomic Claim 193/372 · 2026-03-24_nvidia-the-inference-kingdom-expands::NIEK2-0192::SPLIT02
Atomic Claim 194/372 · 2026-03-24_nvidia-the-inference-kingdom-expands::NIEK2-0193
Nvidia’s approach has been to use copper where they can, and optics where they must. The architecture of NVL1152 in the Feynman generation will follow the same principle. It is clear that the NVL1152 will adopt CPO to connect between racks, but from GPUs to NVLink Switches is currently copper POR. Nvidia is unable to achieve another doubling of electrical lane speed from 224Gbit/s bi-di to 448Gbit/s uni-di means bandwidth isn’t that amazing.
Atomic Claim 195/372 · 2026-03-24_nvidia-the-inference-kingdom-expands::NIEK2-0194
Atomic Claim 196/372 · 2026-03-24_nvidia-the-inference-kingdom-expands::NIEK2-0195
Atomic Claim 197/372 · 2026-03-24_nvidia-the-inference-kingdom-expands::NIEK2-0196
Claim: Feynman generation 的 NVL1152 架構將遵循相同原則。
Frame:COMPARISON· Mode:ASSERTED· Mapping:PARTIAL
開啟逐條審核
Atomic Claim 198/372 · 2026-03-24_nvidia-the-inference-kingdom-expands::NIEK2-0197
Atomic Claim 199/372 · 2026-03-24_nvidia-the-inference-kingdom-expands::NIEK2-0198
Atomic Claim 200/372 · 2026-03-24_nvidia-the-inference-kingdom-expands::NIEK2-0199
While 448G high speed SerDes have big challenges for shoreline, reach, and power versus using a die-to-die connection to an optical engine, the manufacturing challenges, cost, and reliability for Feynman necessitate using copper to the Switch.
Atomic Claim 201/372 · 2026-03-24_nvidia-the-inference-kingdom-expands::NIEK2-0200
Atomic Claim 202/372 · 2026-03-24_nvidia-the-inference-kingdom-expands::NIEK2-0201
Claim: 相較以 die-to-die 方式連接 optical engine,448G 高速 SerDes 在功耗方面也有重大挑戰;考量 Feynman 的製造難度、成本與可靠性,系統仍需要以 copper 連到 Switch。
Frame:COMPARISON· Mode:ASSERTED· Mapping:COMPLETE
開啟逐條審核
With that said, the NVL1152 SKU is years out – and the roadmap is highly likely to shift. For now, our base case stands at copper being used within each rack and CPO between the racks, but this could easily change.
Atomic Claim 203/372 · 2026-03-24_nvidia-the-inference-kingdom-expands::NIEK2-0202
Atomic Claim 204/372 · 2026-03-24_nvidia-the-inference-kingdom-expands::NIEK2-0203
Atomic Claim 205/372 · 2026-03-24_nvidia-the-inference-kingdom-expands::NIEK2-0204
For now – our best estimate of Nvidia’s CPO roadmap is as follows:
Atomic Claim 206/372 · 2026-03-24_nvidia-the-inference-kingdom-expands::NIEK2-0205
Rubin:
NVL72 – Oberon all copper scale up
Atomic Claim 207/372 · 2026-03-24_nvidia-the-inference-kingdom-expands::NIEK2-0206
Atomic Claim 208/372 · 2026-03-24_nvidia-the-inference-kingdom-expands::NIEK2-0207
Claim: 以下為 Rubin Ultra 的規劃。
Frame:ATTRIBUTE· Mode:ESTIMATED· Mapping:PARTIAL
開啟逐條審核
NVL72 – Oberon all copper scale up
NVL144 – Kyber rack all copper scale up
Atomic Claim 209/372 · 2026-03-24_nvidia-the-inference-kingdom-expands::NIEK2-0208
Claim: NVL144 採 Kyber rack,scale-up 全部使用 copper。
Frame:NARY_RELATION· Mode:ESTIMATED· Mapping:COMPLETE
開啟逐條審核
NVL288 – Kyber rack all copper scale up with copper connecting 2 racks together
Atomic Claim 210/372 · 2026-03-24_nvidia-the-inference-kingdom-expands::NIEK2-0209
Claim: NVL288 採 Kyber rack,scale-up 全部使用 copper,並以 copper 將兩個機櫃連接 together。
Frame:NARY_RELATION· Mode:ESTIMATED· Mapping:COMPLETE
開啟逐條審核
NVL576 – 8x Oberon Racks copper scale up within rack and CPO on switch between racks in a two tier all to all topology. This would be low volume for test purposes
Atomic Claim 211/372 · 2026-03-24_nvidia-the-inference-kingdom-expands::NIEK2-0210
Atomic Claim 212/372 · 2026-03-24_nvidia-the-inference-kingdom-expands::NIEK2-0211
Atomic Claim 213/372 · 2026-03-24_nvidia-the-inference-kingdom-expands::NIEK2-0212
NVL72 – Oberon Rack – All Copper
Atomic Claim 214/372 · 2026-03-24_nvidia-the-inference-kingdom-expands::NIEK2-0213
NVL144 – Kyber Rack – All Copper
Atomic Claim 215/372 · 2026-03-24_nvidia-the-inference-kingdom-expands::NIEK2-0214
Claim: Feynman NVL144 採 Kyber Rack,全部使用 Copper。
Frame:NARY_RELATION· Mode:ESTIMATED· Mapping:PARTIAL
開啟逐條審核
NVL1152 – 8xKyber Rack – Copper within rack and CPO on the switch between racks
Atomic Claim 216/372 · 2026-03-24_nvidia-the-inference-kingdom-expands::NIEK2-0215

Source: SemiAnalysis, Nvidia
Oberon and Kyber Updates, Larger World Sizes Introduced, More Networking Updates
Nvidia provided a long-awaited update on its Kyber rack form factor, the latest addition to the lineup after Oberon having first been previewed as a prototype at GTC 2025. As a prototype, the rack architecture has continued to evolve, and we notice some changes. First, each compute blade has densified, with 4x Rubin Ultra GPU and 2x Vera each. There are a total of 2 canisters of 18 compute blades which amounts to 36 compute blades total for 144 GPUs in a rack. The initial Kyber design featured 2 GPUs and 2 Vera CPUs in one compute blade, with a total of 4 canisters of 18 compute blades each.
Atomic Claim 217/372 · 2026-03-24_nvidia-the-inference-kingdom-expands::NIEK2-0216
Claim: Nvidia 終於更新了 Kyber rack form factor;Kyber 是繼 Oberon 之後的新成員,最早於 GTC 2025 以 prototype 形式亮相。
Frame:ATTRIBUTE· Mode:ASSERTED· Mapping:COMPLETE
開啟逐條審核
Atomic Claim 218/372 · 2026-03-24_nvidia-the-inference-kingdom-expands::NIEK2-0219
Claim: 每個 compute blade 配置 4 顆 Rubin Ultra GPU。
Frame:RELATION· Mode:ASSERTED· Mapping:COMPLETE
開啟逐條審核
Atomic Claim 219/372 · 2026-03-24_nvidia-the-inference-kingdom-expands::NIEK2-0221
Atomic Claim 220/372 · 2026-03-24_nvidia-the-inference-kingdom-expands::NIEK2-0222
The details below are based on the Rubin Kyber prototypes, but Rubin Ultra will be redone.
Atomic Claim 221/372 · 2026-03-24_nvidia-the-inference-kingdom-expands::NIEK2-0223
Atomic Claim 222/372 · 2026-03-24_nvidia-the-inference-kingdom-expands::NIEK2-0224
Claim: Rubin Ultra 版本將重新設計。
Frame:ATTRIBUTE· Mode:ASSERTED· Mapping:PARTIAL
開啟逐條審核

Source: SemiAnalysis
Each switch blade is also double in height vs the GTC 2025 prototype, with 6 NVLink 7 switches per switch blade, and 12 switch blades per rack, amounting to a total of 72 NVLink 7 switches per Kyber rack. The GPUs are connected all-to-all to the switch blades via 2 PCB midplanes or 1 midplane per canister.
Atomic Claim 223/372 · 2026-03-24_nvidia-the-inference-kingdom-expands::NIEK2-0225
Claim: 每個 switch blade 的高度也較 GTC 2025 prototype 增加一倍;每個 switch blade 配置 6 顆 NVLink 7 switches,每櫃有 12 個 switch blades,因此一個 Kyber rack 共 72 顆 NVLink 7 switches。
Frame:COMPARISON· Mode:ASSERTED· Mapping:COMPLETE
開啟逐條審核
Atomic Claim 224/372 · 2026-03-24_nvidia-the-inference-kingdom-expands::NIEK2-0226

Kyber midplane PCB (GPU side). Source: Nvidia, SemiAnalysis
Atomic Claim 225/372 · 2026-03-24_nvidia-the-inference-kingdom-expands::NIEK2-0227
For Rubin Ultra NVL144 Kyber, there will be no CPO used for scale up as we have told clients multiple times ↗, despite rumors from other analysts suggesting scale-up CPO introduction for Kyber. However, optics for NVLink are coming and will be progressively phased in. Scale-up CPO will first be used for the Rubin Ultra NVL 576 system to connect between 8 Oberon form factor racks, forming a two-layer all-to-all network. A copper backplane will still be used for scale-up networking within the racks however. This is still for low volume / testing purposes.
Atomic Claim 226/372 · 2026-03-24_nvidia-the-inference-kingdom-expands::NIEK2-0228
Atomic Claim 227/372 · 2026-03-24_nvidia-the-inference-kingdom-expands::NIEK2-0229
Atomic Claim 228/372 · 2026-03-24_nvidia-the-inference-kingdom-expands::NIEK2-0230
Atomic Claim 229/372 · 2026-03-24_nvidia-the-inference-kingdom-expands::NIEK2-0231
Claim: Scale-up CPO 首先會用在 Rubin Ultra NVL576,連接 8 個 Oberon form factor racks,形成兩層 all-to-all network。
Frame:NARY_RELATION· Mode:ASSERTED· Mapping:PARTIAL
開啟逐條審核
Atomic Claim 230/372 · 2026-03-24_nvidia-the-inference-kingdom-expands::NIEK2-0232
Claim: 不過機櫃內的 scale-up networking 仍會使用 copper backplane。
Frame:NARY_RELATION· Mode:ASSERTED· Mapping:PARTIAL
開啟逐條審核
Moving back to the Kyber Rack, each Rubin Ultra logical GPU offers 14.4Tbit/s uni-di of scale-up bandwidth, using an 80DP connector (72 DPs used x 200Gbit/s bi-di channel = 14.4Tbit/s) per GPU for connectivity to the midplane board. Connecting all 144 GPUs in an all-to-all network will require 72 NVLink 7.0 Switch Chips running at 28.8Tbit/s uni-di of aggregate bandwidth each.
Atomic Claim 231/372 · 2026-03-24_nvidia-the-inference-kingdom-expands::NIEK2-0234
Claim: 回到 Kyber Rack,每顆 Rubin Ultra logical GPU 提供 14.4Tbit/s uni-di scale-up 頻寬;每顆 GPU 以一個 80DP connector(使用 72 DPs × 200Gbit/s bi-di channel=14.4Tbit/s)連到 midplane board。
Frame:COMPARISON· Mode:ASSERTED· Mapping:COMPLETE
開啟逐條審核
Atomic Claim 232/372 · 2026-03-24_nvidia-the-inference-kingdom-expands::NIEK2-0235

Source: SemiAnalysis
In the Kyber Switch Blade picture below, we can see that there are 2 separate PCBs carrying 3 Switches each. The switch blade should have 6 152DP connectors, 3 connectors serving each midplane board. The picture is a prototype blade using less dense connectors, which is why there are 12 connectors instead of the 6 that we expect in the production version.
Atomic Claim 233/372 · 2026-03-24_nvidia-the-inference-kingdom-expands::NIEK2-0236
Atomic Claim 234/372 · 2026-03-24_nvidia-the-inference-kingdom-expands::NIEK2-0237

Source: Nvidia, SemiAnalysis
Each 28.8T NVLink Switch has 144 lanes of 200G (simultaneous bi-directional) which means each Switch has 24 lanes of 200G going to each connector. Copper flyover cables are used to connect each switch to the midplane, as the distances involved are too long for PCB traces. This is also why the switches are further away from the midplane, to provide space for the routing of the flyover cables.
Atomic Claim 235/372 · 2026-03-24_nvidia-the-inference-kingdom-expands::NIEK2-0239
Atomic Claim 236/372 · 2026-03-24_nvidia-the-inference-kingdom-expands::NIEK2-0240
Claim: 每顆 switch 與 midplane 之間使用 Copper flyover cables,因為距離太長,不適合直接用 PCB traces。
Frame:NARY_RELATION· Mode:ASSERTED· Mapping:PARTIAL
開啟逐條審核
Atomic Claim 237/372 · 2026-03-24_nvidia-the-inference-kingdom-expands::NIEK2-0241
Claim: 這也是 switches 必須與 midplane 保持較大距離的原因,以騰出 flyover cables routing 空間。
Frame:COMPARISON· Mode:ASSERTED· Mapping:PARTIAL
開啟逐條審核

Source: SemiAnalysis Networking Model ↗
Each NVLink Switch Chip connects via flyover cables to the connector (144 DPs used x 200 Gbit/s bi-di channel = 28.8Tbit/s) connectors at the edge of the switch blade, and these connectors plug into the midplane board. Nvidia is looking into using co-packaged Copper to reduce loss further, in case NPC doesn’t work. As far as we know the Nvidia is telling supply chain to go for fully co-packaged copper.
Atomic Claim 238/372 · 2026-03-24_nvidia-the-inference-kingdom-expands::NIEK2-0242
Claim: 每顆 NVLink Switch chip 透過 flyover cables 連到 switch blade 邊緣 connectors;使用 144 DPs × 200Gbit/s bi-di channel,可提供 28.8Tbit/s,這些 connectors 再插入 midplane board。
Frame:COMPARISON· Mode:ASSERTED· Mapping:COMPLETE
開啟逐條審核
Atomic Claim 239/372 · 2026-03-24_nvidia-the-inference-kingdom-expands::NIEK2-0243
Claim: 若 NPC 方案不可行,Nvidia 正研究採用 co-packaged Copper 以進一步降低訊號損耗。
Frame:RELATION· Mode:HYPOTHETICAL· Mapping:COMPLETE
開啟逐條審核
Atomic Claim 240/372 · 2026-03-24_nvidia-the-inference-kingdom-expands::NIEK2-0244
Claim: 依 SemiAnalysis 所知,Nvidia 正要求供應鏈朝全面 co-packaged copper 方向開發。
Frame:RELATION· Mode:ASSERTED· Mapping:COMPLETE
開啟逐條審核
Rubin Ultra NVL288
Though not officially discussed by Nvidia at GTC 2026, an NVL288 concept has been explored within the supply chain. This would entail two NVL144 Kyber racks placed adjacent to each other, with a rack-to-rack copper backplane used to connect the two racks. One possibility is that all 288 GPUs are connected all to all, but this would require higher radix switches than the current NVLink 7 switches which only offer a maximum radix of 144 ports of 200G.
Atomic Claim 241/372 · 2026-03-24_nvidia-the-inference-kingdom-expands::NIEK2-0245
Atomic Claim 242/372 · 2026-03-24_nvidia-the-inference-kingdom-expands::NIEK2-0246
Atomic Claim 243/372 · 2026-03-24_nvidia-the-inference-kingdom-expands::NIEK2-0247
Atomic Claim 244/372 · 2026-03-24_nvidia-the-inference-kingdom-expands::NIEK2-0248
Atomic Claim 245/372 · 2026-03-24_nvidia-the-inference-kingdom-expands::NIEK2-0249
If Rubin Ultra NVL288 is deployed, each Rubin Ultra GPU will have a scale-up bandwidth of 14.4Tbit/s uni-di, requiring 144 DPs of cables to connect the NVLink 7 switches. 72 DPs per GPU times 288 GPUs means a total of 20,736 additional DPs required to connect this larger world size domain. This entails a lot of cables, so it is an upper bound of how much cable content could be used.
Atomic Claim 246/372 · 2026-03-24_nvidia-the-inference-kingdom-expands::NIEK2-0250
Claim: 若部署 Rubin Ultra NVL288,每顆 Rubin Ultra GPU 會有 14.4Tbit/s uni-di scale-up 頻寬,並需要 144 DPs cables 連到 NVLink 7 switches。
Frame:NARY_RELATION· Mode:HYPOTHETICAL· Mapping:PARTIAL
開啟逐條審核
Atomic Claim 247/372 · 2026-03-24_nvidia-the-inference-kingdom-expands::NIEK2-0251
Atomic Claim 248/372 · 2026-03-24_nvidia-the-inference-kingdom-expands::NIEK2-0252
The radix of the 28.8T NVLink Switch limits the number of GPUs that each switch can connect while still providing for cross-rack connectivity. Either a higher radix switch will have to be used - or there will have to be a degree of oversubscription in this architecture while potentially adopting a dragonfly-like network topology. This would also require fewer DPs worth of copper cables.
Atomic Claim 249/372 · 2026-03-24_nvidia-the-inference-kingdom-expands::NIEK2-0253
Atomic Claim 250/372 · 2026-03-24_nvidia-the-inference-kingdom-expands::NIEK2-0254
Claim: 因此要嘛必須採用更高 radix 的 switch,要嘛架構需要一定程度的 oversubscription,並可能採用類似 dragonfly 的 network topology。
Frame:ATTRIBUTE· Mode:ASSERTED· Mapping:PARTIAL
開啟逐條審核
Atomic Claim 251/372 · 2026-03-24_nvidia-the-inference-kingdom-expands::NIEK2-0255

Source: SemiAnalysis
All current evidence in the supply chain points to NVSwitch 7 being the same bandwidth as NVSwitch 6, but that is seems a bit illogical to be frank. Our belief is that NVSwitch 7 is actually 2x the bandwidth and radix of NVSwitch 6, so all-to-all can be done, and architecturally that makes the most sense from a systems perspective.
Atomic Claim 252/372 · 2026-03-24_nvidia-the-inference-kingdom-expands::NIEK2-0256
Atomic Claim 253/372 · 2026-03-24_nvidia-the-inference-kingdom-expands::NIEK2-0257
Atomic Claim 254/372 · 2026-03-24_nvidia-the-inference-kingdom-expands::NIEK2-0258
Atomic Claim 255/372 · 2026-03-24_nvidia-the-inference-kingdom-expands::NIEK2-0259
Rubin Ultra NVL576
To push the scale up world size beyond 144 GPUs and across multiple racks, optics are needed as we are approaching the maximum compute density that is within the reach of copper. Rubin Ultra NVL576 is now on the roadmap with 8 racks of lower density Oberon.
Atomic Claim 256/372 · 2026-03-24_nvidia-the-inference-kingdom-expands::NIEK2-0260
Atomic Claim 257/372 · 2026-03-24_nvidia-the-inference-kingdom-expands::NIEK2-0261
Claim: Rubin Ultra NVL576 已進入 roadmap,採用 8 個較低密度的 Oberon racks。
Frame:COMPARISON· Mode:ASSERTED· Mapping:COMPLETE
開啟逐條審核

Source: SemiAnalysis
Optics will be required for the inter-rack connections, though strictly speaking it isn’t confirmed whether this will be with pluggable optics or with CPO, though CPO seems much more likely. The current Blackwell NVL576 prototype “Polyphe” uses pluggable optics.
Atomic Claim 258/372 · 2026-03-24_nvidia-the-inference-kingdom-expands::NIEK2-0262
Claim: 機櫃間連線必須使用 optics;目前尚未確認會採 pluggable optics 還是 CPO,但 SemiAnalysis 認為 CPO 的可能性高得多。
Frame:RELATION· Mode:ASSERTED· Mapping:PARTIAL
開啟逐條審核
Atomic Claim 259/372 · 2026-03-24_nvidia-the-inference-kingdom-expands::NIEK2-0263
Claim: 目前的 Blackwell NVL576 prototype「Polyphe」採用 pluggable optics。
Frame:NARY_RELATION· Mode:ASSERTED· Mapping:COMPLETE
開啟逐條審核
We have shown a concept of NVL576 for GB200 previously ↗ with pluggable optics to interconnect the second layer of NVLink switches. The use of pluggables contributed to an enormous increase in BOM cost that made the system untenable from a TCO perspective for a switched all-to-all. However, it is plausible that Rubin Ultra NVL576 will be rolled out in test volumes before Feynman NVL 1,152, where we will see actual volume ramp of scale-up CPO.
Atomic Claim 260/372 · 2026-03-24_nvidia-the-inference-kingdom-expands::NIEK2-0264
Claim: SemiAnalysis 先前曾展示 GB200 NVL576 概念,以 pluggable optics 互連第二層 NVLink switches。
Frame:ATTRIBUTE· Mode:ASSERTED· Mapping:PARTIAL
開啟逐條審核
Atomic Claim 261/372 · 2026-03-24_nvidia-the-inference-kingdom-expands::NIEK2-0265
Atomic Claim 262/372 · 2026-03-24_nvidia-the-inference-kingdom-expands::NIEK2-0266
The downstream implications of this are exposed in our institutional research, trusted by all major hyperscalers, semiconductor companies, and AI Labs, at sales@semianalysis.com
Feynman
While not much is known about Feynman, the Keynote sneak peek was enough to tell us Feynman will be exciting, with three major technical innovations all being pushed in a single platform: Hybrid bonding/SoIC ↗, A16, CPO ↗, and custom HBM ↗.
Atomic Claim 263/372 · 2026-03-24_nvidia-the-inference-kingdom-expands::NIEK2-0267
Claim: Feynman 將採用 Hybrid bonding/SoIC。
Frame:NARY_RELATION· Mode:ASSERTED· Mapping:COMPLETE
開啟逐條審核
Atomic Claim 264/372 · 2026-03-24_nvidia-the-inference-kingdom-expands::NIEK2-0268
Atomic Claim 265/372 · 2026-03-24_nvidia-the-inference-kingdom-expands::NIEK2-0269
Atomic Claim 266/372 · 2026-03-24_nvidia-the-inference-kingdom-expands::NIEK2-0270
While Feynman adopting CPO is on the roadmap, the question is to what extent? Will in-rack interconnectivity be copper based or optical? We will show possible configurations behind the Paywall. Vera ETL256
Atomic Claim 267/372 · 2026-03-24_nvidia-the-inference-kingdom-expands::NIEK2-0271
Atomic Claim 268/372 · 2026-03-24_nvidia-the-inference-kingdom-expands::NIEK2-0272
Atomic Claim 269/372 · 2026-03-24_nvidia-the-inference-kingdom-expands::NIEK2-0273
Claim: 此處主題為 Vera ETL256。
Frame:ATTRIBUTE· Mode:ASSERTED· Mapping:PARTIAL
開啟逐條審核
CPU demand is rising as AI workloads require more data handling, preprocessing, and orchestration beyond GPU compute. Reinforcement learning further increases demand, with CPUs running simulations, executing code, and verifying outputs in parallel. As GPUs scale faster than CPUs, larger CPU clusters are needed to keep them fully utilized, making CPUs a growing bottleneck.
Atomic Claim 270/372 · 2026-03-24_nvidia-the-inference-kingdom-expands::NIEK2-0274
Atomic Claim 271/372 · 2026-03-24_nvidia-the-inference-kingdom-expands::NIEK2-0275
Claim: Reinforcement learning 會進一步推升 CPU 需求。
Frame:ATTRIBUTE· Mode:ASSERTED· Mapping:PARTIAL
開啟逐條審核
Atomic Claim 272/372 · 2026-03-24_nvidia-the-inference-kingdom-expands::NIEK2-0276
Atomic Claim 273/372 · 2026-03-24_nvidia-the-inference-kingdom-expands::NIEK2-0277
The Vera standalone rack addresses this directly, achieving unprecedented density by fitting 256 CPUs into a single rack — a feat that necessitates liquid cooling. The underlying rationale mirrors the NVL rack design philosophy: pack compute tightly enough that copper interconnects can reach everything within the rack, eliminating the need for optical transceivers on the spine. The cost savings from copper more than offset the additional cooling overhead.
Atomic Claim 274/372 · 2026-03-24_nvidia-the-inference-kingdom-expands::NIEK2-0278
Claim: Vera standalone rack 直接針對此問題設計,單一機櫃可容納 256 顆 CPUs,達到前所未有的密度,也因此必須使用 liquid cooling。
Frame:RELATION· Mode:ASSERTED· Mapping:PARTIAL
開啟逐條審核
Atomic Claim 275/372 · 2026-03-24_nvidia-the-inference-kingdom-expands::NIEK2-0279
Claim: 其設計邏輯與 NVL rack 相同:把運算元件塞得足夠密集,使 copper interconnect 能涵蓋整個機櫃,因而不需要在 spine 使用 optical transceivers。
Frame:COMPARISON· Mode:ASSERTED· Mapping:COMPLETE
開啟逐條審核
Atomic Claim 276/372 · 2026-03-24_nvidia-the-inference-kingdom-expands::NIEK2-0280

Source: SemiAnalysis
Each Vera ETL rack consists of 32 compute trays, 16 above and 16 below, arranged symmetrically around four 1U MGX ETL switch trays (based on Spectrum-6) in the middle. The symmetric split is deliberate: it minimizes cable length variance between compute trays and the spine, keeping all connections within copper reach. From each switch tray, rear-facing ports connect to that copper spine for intra-rack communication, while 32 front-facing OSFP cages provide optical connectivity to the rest of the POD.
Atomic Claim 277/372 · 2026-03-24_nvidia-the-inference-kingdom-expands::NIEK2-0283
Atomic Claim 278/372 · 2026-03-24_nvidia-the-inference-kingdom-expands::NIEK2-0281
Atomic Claim 279/372 · 2026-03-24_nvidia-the-inference-kingdom-expands::NIEK2-0282
Atomic Claim 280/372 · 2026-03-24_nvidia-the-inference-kingdom-expands::NIEK2-0284
Claim: 上下兩組 compute trays 對稱排列在中央 4 個 1U MGX ETL switch trays 周圍,而這些 switch trays 基於 Spectrum-6。
Frame:ATTRIBUTE· Mode:ASSERTED· Mapping:COMPLETE
開啟逐條審核
Atomic Claim 281/372 · 2026-03-24_nvidia-the-inference-kingdom-expands::NIEK2-0285
Atomic Claim 282/372 · 2026-03-24_nvidia-the-inference-kingdom-expands::NIEK2-0286
Atomic Claim 283/372 · 2026-03-24_nvidia-the-inference-kingdom-expands::NIEK2-0287
Networking within the rack uses a Spectrum-X multiplane topology, distributing 200 Gb/s lanes across the four switches to achieve full all-to-all connectivity while maintaining a single network tier. With each compute tray housing 8 Vera CPUs, the result is 256 CPUs per rack, all interconnected over Ethernet through a single, flat network.
Atomic Claim 284/372 · 2026-03-24_nvidia-the-inference-kingdom-expands::NIEK2-0288
Atomic Claim 285/372 · 2026-03-24_nvidia-the-inference-kingdom-expands::NIEK2-0289
Claim: 每個 compute tray 配置 8 顆 Vera CPUs,因此每櫃共 256 顆 CPUs,並透過單一 flat network 以 Ethernet 全部互連。
Frame:NARY_RELATION· Mode:ASSERTED· Mapping:COMPLETE
開啟逐條審核

Source: SemiAnalysis

CMX and STX
We have written extensively on Nvidia’s CMX, or ICMS platform in our last Rubin piece and Memory Model. Nvidia introduced the STX reference storage rack architecture.
Atomic Claim 286/372 · 2026-03-24_nvidia-the-inference-kingdom-expands::NIEK2-0290
Claim: Nvidia 推出了 STX reference storage rack 架構。
Frame:RELATION· Mode:ASSERTED· Mapping:COMPLETE
開啟逐條審核
CMX
CMX is NVIDIA’s context memory storage platform. CMX addresses a growing bottleneck in modern inference infrastructure: the rapid expansion of KV Cache required to support long-context and agentic workloads.
KV cache grows linearly with input sequence length and number of users and is the primary tradeoff when it comes to prefill performance (time to first token). At scale, on-device HBM does not have enough capacity. Host DRAM extends beyond HBM capacity with an additional tier of cache, but also hits limits on total amount per node, memory bandwidth, and network bandwidth. Enter NVMe storage for additional KVcache offload.
Atomic Claim 287/372 · 2026-03-24_nvidia-the-inference-kingdom-expands::NIEK2-0294
Claim: 在 prefill 效能(time to first token)方面,KV cache 是主要 trade-off。
Frame:ATTRIBUTE· Mode:ASSERTED· Mapping:COMPLETE
開啟逐條審核
Atomic Claim 288/372 · 2026-03-24_nvidia-the-inference-kingdom-expands::NIEK2-0293
Claim: KV cache 會隨輸入 sequence length 與使用者數量線性成長。
Frame:RELATION· Mode:ASSERTED· Mapping:COMPLETE
開啟逐條審核
Atomic Claim 289/372 · 2026-03-24_nvidia-the-inference-kingdom-expands::NIEK2-0295
Atomic Claim 290/372 · 2026-03-24_nvidia-the-inference-kingdom-expands::NIEK2-0296
Atomic Claim 291/372 · 2026-03-24_nvidia-the-inference-kingdom-expands::NIEK2-0297
Atomic Claim 292/372 · 2026-03-24_nvidia-the-inference-kingdom-expands::NIEK2-0298
NVIDIA introduced a “new” intermediate storage “tier G3.5” within the inference memory hierarchy at CES in January. Tier G3.5 NVMe sits in between tier G3 DRAM and tier G4 shared storage (also NVMe, or SATA/SAS SSD, or HDD). Previously referred to as ICMS (Inference Context Memory Storage) and now branded as the CMX platform, this is just another re-brand of storage servers attached to compute servers via Bluefield NICs. The only difference from NVMe architectures is the swap from Connect-X NICs to Bluefield NICs.
Atomic Claim 293/372 · 2026-03-24_nvidia-the-inference-kingdom-expands::NIEK2-0299
Atomic Claim 294/372 · 2026-03-24_nvidia-the-inference-kingdom-expands::NIEK2-0301
Claim: 這個平台過去稱為 ICMS(Inference Context Memory Storage),現在則改名為 CMX;本質上就是透過 Bluefield NICs 將 storage servers 連接到 compute servers 的另一種品牌包裝。
Frame:RELATION· Mode:ASSERTED· Mapping:COMPLETE
開啟逐條審核
Atomic Claim 295/372 · 2026-03-24_nvidia-the-inference-kingdom-expands::NIEK2-0300
Atomic Claim 296/372 · 2026-03-24_nvidia-the-inference-kingdom-expands::NIEK2-0302

Source: Original NVIDIA ICMS blog in January, 2026 – updated and re-released on March 16, 2026 https://developer.[[02_companies/NVDA|nvidia]].com/blog/introducing-[[02_companies/NVDA|nvidia]]-[[04_knowledge_base/BlueField-4|bluefield-4]]-powered-inference-context-memory-storage-platform-for-the-next-frontier-of-ai/ ↗
STX
To expand the scope of CMX, NVIDIA also launched STX. STX is a reference rack architecture using Nvidia’s BF-4 based storage solution to complement VR compute racks. The reference architecture effectively specifies exactly how many drives, Vera CPUs, BF-4 DPUs, CX-9 NICs, and Spectrum-X switches are needed for a given cluster.
Atomic Claim 297/372 · 2026-03-24_nvidia-the-inference-kingdom-expands::NIEK2-0304
Atomic Claim 298/372 · 2026-03-24_nvidia-the-inference-kingdom-expands::NIEK2-0305
Atomic Claim 299/372 · 2026-03-24_nvidia-the-inference-kingdom-expands::NIEK2-0306

BF-4 in STX. Source: Nvidia, SemiAnalysis
Atomic Claim 300/372 · 2026-03-24_nvidia-the-inference-kingdom-expands::NIEK2-0307
Unlike the BF-4 in the VR NVL72, which consists of a Grace CPU and a single CX-9 NIC, the BF-4 in the STX reference design includes one Vera CPU, two CX-9 NICs, and two SOCAMM modules. Each STX box contains two BF-4 units, totaling two Vera CPUs, four CX-9 NICs, and four SOCAMM modules. For the whole STX rack, it has a total of 16 boxes, implying 32 Vera CPUs, 64 CX-9 NICs, and 64 SOCAMMs.
Atomic Claim 301/372 · 2026-03-24_nvidia-the-inference-kingdom-expands::NIEK2-0308
Atomic Claim 302/372 · 2026-03-24_nvidia-the-inference-kingdom-expands::NIEK2-0311
Atomic Claim 303/372 · 2026-03-24_nvidia-the-inference-kingdom-expands::NIEK2-0310
Atomic Claim 304/372 · 2026-03-24_nvidia-the-inference-kingdom-expands::NIEK2-0312
Atomic Claim 305/372 · 2026-03-24_nvidia-the-inference-kingdom-expands::NIEK2-0309
Atomic Claim 306/372 · 2026-03-24_nvidia-the-inference-kingdom-expands::NIEK2-0315
Atomic Claim 307/372 · 2026-03-24_nvidia-the-inference-kingdom-expands::NIEK2-0316
Atomic Claim 308/372 · 2026-03-24_nvidia-the-inference-kingdom-expands::NIEK2-0313
Atomic Claim 309/372 · 2026-03-24_nvidia-the-inference-kingdom-expands::NIEK2-0314

STX Rack (left). Source: Nvidia, SemiAnalysis
Atomic Claim 310/372 · 2026-03-24_nvidia-the-inference-kingdom-expands::NIEK2-0317
The STX announcement included a typical Nvidia show of strength where they named all major storage vendors as supporting STX, including AIC, Cloudian, DDN, Dell Technologies, Everpure, Hitachi Vantara, HPE, IBM, MinIO, NetApp, Nutanix, Supermicro, Quanta Cloud Technology (QCT), VAST Data and WEKA.
Atomic Claim 311/372 · 2026-03-24_nvidia-the-inference-kingdom-expands::NIEK2-0318
Put together, BlueField-4, CMX, and STX represent NVIDIA’s broader effort to standardize how clusters are designed at the storage layer. NVIDIA has captured the compute and network layer, and is actively moving into the storage, software, and infrastructure operations layers over time.
Atomic Claim 312/372 · 2026-03-24_nvidia-the-inference-kingdom-expands::NIEK2-0320
Atomic Claim 313/372 · 2026-03-24_nvidia-the-inference-kingdom-expands::NIEK2-0322
Atomic Claim 314/372 · 2026-03-24_nvidia-the-inference-kingdom-expands::NIEK2-0321
Now behind the paywall, we will share some more details on how all of this impacts the supply chain. Including beneficiaries of the LPX system, and the updated Kyber racks. We will also reveal a rack concept that Nvidia has yet to announce.
Atomic Claim 315/372 · 2026-03-24_nvidia-the-inference-kingdom-expands::NIEK2-0323
Atomic Claim 316/372 · 2026-03-24_nvidia-the-inference-kingdom-expands::NIEK2-0324
Feynman NVL1152 Networking Topologies
Within each Feynman Kyber rack, we tentatively assume double the bandwidth per logical GPU and double the NVLink Switch bandwidth to 28.8T and 57.6T respectively. Though Jensen, in the Financial Q+A the day after the GTC Keynote, characterized NVL1152 as “all CPO”, the key technical blog outlining the new rack form factors ↗ only strictly referenced CPO for rack to rack interconnect. We will discuss the potential topography for both options.
Atomic Claim 317/372 · 2026-03-24_nvidia-the-inference-kingdom-expands::NIEK2-0325
Claim: 在每個 Feynman Kyber rack 中,SemiAnalysis 暫時假設每顆 logical GPU 的頻寬加倍至 28.8T,而 NVLink Switch 頻寬加倍至 57.6T。
Frame:COMPARISON· Mode:HYPOTHETICAL· Mapping:COMPLETE
開啟逐條審核
Atomic Claim 318/372 · 2026-03-24_nvidia-the-inference-kingdom-expands::NIEK2-0326
To double the scale-up bandwidth using copper interconnect, NVIDIA would have to achieve a per lane bandwidth of 448Gbit/s uni-di (and implemented with simulatanous bi-directional SerDes so that each physical channel carries 448G of RX and 448G of TX) . However, this is a challenging feat as they would first have to prove the feasibility of 448Gb/s PAM4 SerDes at large volumes, then implement echo cancellation to achieve bidirectional bandwidth ↗, which is in itself extremely difficult. We believe Nvidia is going for 448G uni-di only.
Atomic Claim 319/372 · 2026-03-24_nvidia-the-inference-kingdom-expands::NIEK2-0328
Claim: 但這非常困難,因為首先必須證明 448Gb/s PAM4 SerDes 可大規模量產,再導入 echo cancellation 以實現雙向頻寬,而後者本身也極具挑戰。
Frame:NARY_RELATION· Mode:ASSERTED· Mapping:PARTIAL
開啟逐條審核
Atomic Claim 320/372 · 2026-03-24_nvidia-the-inference-kingdom-expands::NIEK2-0327
Atomic Claim 321/372 · 2026-03-24_nvidia-the-inference-kingdom-expands::NIEK2-0329

Source: SemiAnalysis
Feynman could use in-rack optics, where switch blades are blind-mated to the midplane using optical connectors and thin fiber strands can be used to connect the optical connectors to the NVLink 8 Switches in place of flyover cables., but we believe this is very unlikely.
Atomic Claim 322/372 · 2026-03-24_nvidia-the-inference-kingdom-expands::NIEK2-0330
Atomic Claim 323/372 · 2026-03-24_nvidia-the-inference-kingdom-expands::NIEK2-0331

Source: SemiAnalysis
For rack-to-rack interconnect, we explore two different topologies. The first is a two-layer CLOS network that is similar to the Oberon form factor, but with twice the bandwidth of each GPU and NVLink switch.
Atomic Claim 324/372 · 2026-03-24_nvidia-the-inference-kingdom-expands::NIEK2-0332
Atomic Claim 325/372 · 2026-03-24_nvidia-the-inference-kingdom-expands::NIEK2-0333
Claim: 第一種是兩層 CLOS network,與 Oberon form factor 類似,但每顆 GPU 與 NVLink switch 的頻寬都加倍。
Frame:COMPARISON· Mode:HYPOTHETICAL· Mapping:COMPLETE
開啟逐條審核

Source: SemiAnalysis
The second is a reconfigurable dragonfly topology using OCS switches to connect the 8 racks. The number of OCS ports required for this topology remains tentative.
Atomic Claim 326/372 · 2026-03-24_nvidia-the-inference-kingdom-expands::NIEK2-0334
Atomic Claim 327/372 · 2026-03-24_nvidia-the-inference-kingdom-expands::NIEK2-0335

Source: SemiAnalysis
GTC 2026 Supply Chain Implications
Here, we will discuss our findings on where we see there are big changes in content for the supply chain coming out of all these announcements at GTC.
AlphaWave 112G Serdes in LP30
It may surprise readers that Qualcomm has IP in the Groq LPU 3 chip! More specifically it is AlphaWave, which Qualcomm acquired last year, that is providing the 112G SerDes for Groq’s C2C. AlphaWave was selected as the only IP provider that has high speed SerDes for Samsung Foundry. It was AlphaWave’s SerDes that caused issues for Groq LPU 2. Alphawave will continue to be used for the LP35, but Nvidia will of course use their own NVLink SerDes IP from LP40 when it transitions back to TSMC.
Atomic Claim 328/372 · 2026-03-24_nvidia-the-inference-kingdom-expands::NIEK2-0336
Atomic Claim 329/372 · 2026-03-24_nvidia-the-inference-kingdom-expands::NIEK2-0337
Atomic Claim 330/372 · 2026-03-24_nvidia-the-inference-kingdom-expands::NIEK2-0338
Atomic Claim 331/372 · 2026-03-24_nvidia-the-inference-kingdom-expands::NIEK2-0339
Atomic Claim 332/372 · 2026-03-24_nvidia-the-inference-kingdom-expands::NIEK2-0340
Atomic Claim 333/372 · 2026-03-24_nvidia-the-inference-kingdom-expands::NIEK2-0341
LPX PCB
Next, we mentioned that a very high spec PCB is required for the LPX compute tray. We estimate that each compute tray main board PCB will carry $7k ASP. The suppliers for this are Victory Giant and WUS. Of course, there are several other PCB modules in the compute tray, but they do not need a high spec. Nvidia is continuing with their cable-less philosophy similar to the Vera Rubin compute tray which requires a lot of board-to-board connectors, which brings us to the next big beneficiary.
Atomic Claim 334/372 · 2026-03-24_nvidia-the-inference-kingdom-expands::NIEK2-0342
Claim: 前文提到,LPX compute tray 需要非常高規格的 PCB。
Frame:ATTRIBUTE· Mode:ASSERTED· Mapping:COMPLETE
開啟逐條審核
Atomic Claim 335/372 · 2026-03-24_nvidia-the-inference-kingdom-expands::NIEK2-0343
Claim: SemiAnalysis 估計每片 compute tray main board PCB 的 ASP 約為 7,000 美元。
Frame:ATTRIBUTE· Mode:ESTIMATED· Mapping:COMPLETE
開啟逐條審核
Atomic Claim 336/372 · 2026-03-24_nvidia-the-inference-kingdom-expands::NIEK2-0344
Atomic Claim 337/372 · 2026-03-24_nvidia-the-inference-kingdom-expands::NIEK2-0345
Claim: compute tray 內另外還有數個其他 PCB modules。
Frame:ATTRIBUTE· Mode:ASSERTED· Mapping:COMPLETE
開啟逐條審核
Atomic Claim 338/372 · 2026-03-24_nvidia-the-inference-kingdom-expands::NIEK2-0346
Atomic Claim 339/372 · 2026-03-24_nvidia-the-inference-kingdom-expands::NIEK2-0347
Claim: Nvidia 延續與 Vera Rubin compute tray 類似的 cable-less 設計理念,因此需要大量 board-to-board connectors。
Frame:NARY_RELATION· Mode:ASSERTED· Mapping:COMPLETE
開啟逐條審核
Cables and Connectors: Amphenol Continues to Benefit
For the LPX, Amphenol will be a beneficiary for all the connectors for the backplane. Each LPX node requires 16 80DP Paladin connectors for the backplane. There are also board to board connectors required to connect all the various modules within the tray: the main LPU board with the host CPU module and the OSFP/QSFP modules that sit below the CPU module, the front-end NIC module, and the management module. Amphenol will supply the cable backplane too which is 8,160 DP per rack.
Atomic Claim 340/372 · 2026-03-24_nvidia-the-inference-kingdom-expands::NIEK2-0348
Atomic Claim 341/372 · 2026-03-24_nvidia-the-inference-kingdom-expands::NIEK2-0349
Atomic Claim 342/372 · 2026-03-24_nvidia-the-inference-kingdom-expands::NIEK2-0350
Atomic Claim 343/372 · 2026-03-24_nvidia-the-inference-kingdom-expands::NIEK2-0351
NVL288 System
For the Vera Rubin Ultra NVL288 System that we discussed above, we would say the cable backplane return for Kyber. If Rubin Ultra is deployed in such a form factor – each of the Rubin Ultra GPUs will have a scale-up bandwidth of 14.4Tbit/s uni-di, requiring 144 DPs of cables to connect to the NVSwitches. 144 DPs times 288 GPUs means a total of 41,472 DPs to connect this larger world size domain. This is a lot of cables, so it is more of an upper bound of how much cable content could be used here. If there is oversubscription or if the inter-rack connection is made through the switches – it is possible fewer DPs would be needed.
Atomic Claim 344/372 · 2026-03-24_nvidia-the-inference-kingdom-expands::NIEK2-0352
Claim: 對前述 Vera Rubin Ultra NVL288 系統而言,SemiAnalysis 認為 cable backplane 會重新回到 Kyber。
Frame:ATTRIBUTE· Mode:INFERRED· Mapping:PARTIAL
開啟逐條審核
Atomic Claim 345/372 · 2026-03-24_nvidia-the-inference-kingdom-expands::NIEK2-0353
Claim: 若 Rubin Ultra 採此 form factor,每顆 Rubin Ultra GPUs 都會提供 14.4Tbit/s uni-di scale-up 頻寬,並需要 144 DPs cables 連到 NVSwitches。
Frame:NARY_RELATION· Mode:HYPOTHETICAL· Mapping:PARTIAL
開啟逐條審核
Atomic Claim 346/372 · 2026-03-24_nvidia-the-inference-kingdom-expands::NIEK2-0354
Atomic Claim 347/372 · 2026-03-24_nvidia-the-inference-kingdom-expands::NIEK2-0355
Atomic Claim 348/372 · 2026-03-24_nvidia-the-inference-kingdom-expands::NIEK2-0356
FIT Joining the Party
Backplane cable cartridge and Paladin connector demand is so strong that Amphenol cannot keep up with supply. Amphenol has now completed licensing of the VR NVL72 backplane cable cartridge as well as Paladin HD connectors to FIT, who can now manufacture these components. This has been in the works for a long time but is finally settled. Amphenol will earn licensing fees from FIT’s sales of these licensed components.
Atomic Claim 349/372 · 2026-03-24_nvidia-the-inference-kingdom-expands::NIEK2-0357
Claim: Backplane cable cartridge 與 Paladin connector 需求強到 Amphenol 已無法單獨滿足供應。
Frame:NARY_RELATION· Mode:ASSERTED· Mapping:PARTIAL
開啟逐條審核
Atomic Claim 350/372 · 2026-03-24_nvidia-the-inference-kingdom-expands::NIEK2-0358
Atomic Claim 351/372 · 2026-03-24_nvidia-the-inference-kingdom-expands::NIEK2-0360
Kyber Voronoi – Another FIT Win?
The Kyber midplane will utilize many 8×19 DP connectors to interface with the compute trays at the front of the rack, and to the switch blades in the back of the rack.
Atomic Claim 352/372 · 2026-03-24_nvidia-the-inference-kingdom-expands::NIEK2-0361
For Kyber, Nvidia is now in the driver’s seat when it comes to IP and they have designed a proprietary connector spec named Voronoi, so it will no longer be the Amphenol Paladin connector. There are three vendors bidding for the project: FIT, Molex and Amphenol. FIT appears to be leading the market for these connectors, but Amphenol is reportedly also working together closely with FIT to manufacture the connectors. The design and implementation of Voronoi remains in flux, but both FIT and Amphenol will need to ramp significant production volume with the specification licensed from Nvidia.
Atomic Claim 353/372 · 2026-03-24_nvidia-the-inference-kingdom-expands::NIEK2-0362
Atomic Claim 354/372 · 2026-03-24_nvidia-the-inference-kingdom-expands::NIEK2-0363
Atomic Claim 355/372 · 2026-03-24_nvidia-the-inference-kingdom-expands::NIEK2-0364
Atomic Claim 356/372 · 2026-03-24_nvidia-the-inference-kingdom-expands::NIEK2-0365
Atomic Claim 357/372 · 2026-03-24_nvidia-the-inference-kingdom-expands::NIEK2-0366
Atomic Claim 358/372 · 2026-03-24_nvidia-the-inference-kingdom-expands::NIEK2-0367
The midplane, switch tray and compute tray all feature female side connectors which will require the use of a spring-loaded male part that protect the pins and interface between both sides. The density of these connectors will ultimately be much higher than Amphenol’s Paladin connectors.
Atomic Claim 359/372 · 2026-03-24_nvidia-the-inference-kingdom-expands::NIEK2-0368
Claim: midplane、switch tray 與 compute tray 都使用 female-side connectors,因此需要一個具彈簧機構的 male part 保護 pins,並作為兩側之間的介接元件。
Frame:RELATION· Mode:ASSERTED· Mapping:PARTIAL
開啟逐條審核
Atomic Claim 360/372 · 2026-03-24_nvidia-the-inference-kingdom-expands::NIEK2-0369
More is available to our institutional subscribers sales@semianalysis.com.
Mid-board Optics – Nvidia’s War on Pluggables
Interestingly, the Kyber rack exhibited at the GTC 2026 show floor is missing OSFP cages for scale-out networking. Instead, we only see 4x MPO ports from each compute tray. This design has effectively taken key pluggable transceiver items (driver, TIA, etc.) other than the DSP and put them on a Midboard Optical Module (MBOM) which then connects to the PCB via a land grid array (LGA) socket. Two CX-9s share one MBOM, which then connects to the MPO faceplate via a short fiber connection. The MBOM provides two MPO ports at 2x800G each for 1.6T of total connectivity.
Atomic Claim 361/372 · 2026-03-24_nvidia-the-inference-kingdom-expands::NIEK2-0370
Claim: 值得注意的是,GTC 2026 展場上的 Kyber rack 沒有用於 scale-out networking 的 OSFP cages。
Frame:ATTRIBUTE· Mode:ASSERTED· Mapping:COMPLETE
開啟逐條審核
Atomic Claim 362/372 · 2026-03-24_nvidia-the-inference-kingdom-expands::NIEK2-0371
Claim: 取而代之的是每個 compute tray 只有 4 個 MPO ports。
Frame:ATTRIBUTE· Mode:ASSERTED· Mapping:COMPLETE
開啟逐條審核
Atomic Claim 363/372 · 2026-03-24_nvidia-the-inference-kingdom-expands::NIEK2-0372
Claim: 這個設計等於把 pluggable transceiver 中的關鍵元件,例如 driver、TIA 等移出傳統 pluggable module。
Frame:RELATION· Mode:ASSERTED· Mapping:COMPLETE
開啟逐條審核
Atomic Claim 364/372 · 2026-03-24_nvidia-the-inference-kingdom-expands::NIEK2-0373
Claim: 除了 DSP 外,其他元件被整合到 Midboard Optical Module(MBOM)上,再透過 land grid array(LGA)socket 與 PCB 連接。
Frame:NARY_RELATION· Mode:ASSERTED· Mapping:COMPLETE
開啟逐條審核
Atomic Claim 365/372 · 2026-03-24_nvidia-the-inference-kingdom-expands::NIEK2-0374
Atomic Claim 366/372 · 2026-03-24_nvidia-the-inference-kingdom-expands::NIEK2-0375

4 MPO ports on the left, rather than OSFP cages. Source: Nvidia, SemiAnalysis
Atomic Claim 367/372 · 2026-03-24_nvidia-the-inference-kingdom-expands::NIEK2-0376
The use of MBOM would block the use of any form of pluggable transceiver or AEC, and naturally hyperscalers are saying “CP-Hell No” to that idea and are continuing to push for an OSFP cage so they can continue using pluggables.
Atomic Claim 368/372 · 2026-03-24_nvidia-the-inference-kingdom-expands::NIEK2-0377
Claim: 採用 MBOM 後,任何形式的 pluggable transceiver 或 AEC 都無法使用。
Frame:ATTRIBUTE· Mode:ASSERTED· Mapping:COMPLETE
開啟逐條審核
Atomic Claim 369/372 · 2026-03-24_nvidia-the-inference-kingdom-expands::NIEK2-0379
Claim: hyperscalers 仍持續要求保留 OSFP cage,以繼續使用 pluggables。
Frame:RELATION· Mode:ASSERTED· Mapping:COMPLETE
開啟逐條審核
Atomic Claim 370/372 · 2026-03-24_nvidia-the-inference-kingdom-expands::NIEK2-0378
Claim: 因此 hyperscalers 對此方案的反應非常負面。
Frame:ATTRIBUTE· Mode:ASSERTED· Mapping:COMPLETE
開啟逐條審核
It is important to point out that many aspects of the Kyber design are still in flux and there could still be a number of design changes before Kyber racks are actually deployed. After all – the change from the four canister design to a two compute tray canister + one switch blade bank is already a huge change.
Atomic Claim 371/372 · 2026-03-24_nvidia-the-inference-kingdom-expands::NIEK2-0380
Atomic Claim 372/372 · 2026-03-24_nvidia-the-inference-kingdom-expands::NIEK2-0381
Claim: 畢竟,設計已經從四個 canisters 改成兩個 compute tray canisters 加上一組 switch blade bank,這本身就是非常大的變動。
Frame:ATTRIBUTE· Mode:ASSERTED· Mapping:COMPLETE
開啟逐條審核