NIEK2-0157

① SA Source

Context Before

Second, the LPUs also traverse through the FPGAs to reach the host CPU, with the FPGAs converting C2C to PCIe to the CPU.

Third, the FPGAs are connected to the backplane to talk to other FPGAs in the node, we believe this is to help manage control flow and timing of all the LPUs. The FPGAs also bring extra system DRAM of up to 256GB each. This pool of memory can be used for KVCache if the user wants the entire decode process served by the LPX.

Evidence

there will be 2 cages (likely QSFP-DD) that goes to the Spectrum-switches that is used to connect the LPUs and the GPUs for the disaggregated decode system

Context After

LPU Network

The LPU network can be divided into the scale-up ‘C2C’ network and scale-out network which interacts with the Nvidia GPUs through Spectrum-X. First let’s discuss the scale-up network which can be divided into 3 portions: intra-node, inter-node/intra-rack, inter-rack. For C2C within the rack Nvidia announced a total of 640TB/s of scale up bandwidth per rack which comes from 256 LPUs x 90 lanes x 112Gbps/8 x 2 directions = 645TB/s. Note that Nvidia uses the total 112G line rate rather than 100G of effective data rate.

② Atomic Claim

另有 2 個 cages(很可能是 QSFP-DD)連到 Spectrum switches,用來在 disaggregated decode 系統中連接 LPU 與 GPUs

  • Epistemic Mode: EXPECTED
  • Mapping Status: PARTIAL

③ Semantic Frame

{
  "frame_type": "NARY_RELATION",
  "participants": [
    {
      "node": {
        "id": "04_knowledge_base/QSFP",
        "label": "QSFP"
      },
      "role": "endpoint_or_system"
    },
    {
      "node": {
        "id": "04_knowledge_base/Disaggregated Decode",
        "label": "disaggregated decode"
      },
      "role": "connected_entity"
    },
    {
      "node": {
        "id": "04_knowledge_base/GPU",
        "label": "GPUs"
      },
      "role": "connected_entity"
    }
  ],
  "qualifiers": {
    "condition_text": null,
    "numeric_mentions": [
      "2"
    ],
    "temporal_mentions": []
  },
  "relation_type": "CONNECTS_TO"
}

④ Canonical Entity Mapping

RoleSurface LabelCanonical Target
endpoint_or_systemQSFPQSFP
connected_entitydisaggregated decode04_knowledge_base/Disaggregated Decode
connected_entityGPUsGPU

⑤ Human Review

請在 Properties 逐項確認:

  • 原文 → Atomic Claim 是否忠實
  • Atomic Claim → Semantic Frame 是否忠實
  • Canonical Entity mapping 是否正確
  • Epistemic mode 是否保留原文語氣
  • 最後選擇 review_action

Review state

Markdown 內文不是正式 approval。只有 Apply bridge 寫入的 Decision Ledger event 才是正式決策。