IX2-0184

① SA Source

Context Before

As we mention later in the article, this is best understood as _baseline _data and not completely representative of real-world inference, mainly because InferenceX benchmarks on random data and disables prefix caching. In other words, performance/cost will be _at least _this good. It is also important to note that there are not data points for each GPU at _each _interactivity level. Thus we cannot make _exact _comparisons at each degree of interactivity. We nevertheless think the bar chart comparisons presented below are (very) reasonable interpolations in lieu of using exact data points.

Comparing disagg+wideEP configs at this interactivity level, we see just how effective distributed inference techniques are when it comes to both perf/TCO and overall throughput. We also see how large scale up domains (like GB300 and GB200 NVL72) absolutely dominate in total throughput per GPU.

Evidence

It is interesting to note that at this interactivity level (on an 8k1k workload type), the B200 can achieve the best perf/TCO when MTP is enabled

Context After

image

Source: SemiAnalysis TCO Model

② Atomic Claim

在此 interactivity 水準(8k1k workload)下,開啟 MTPB200 可以達到最佳 perf/TCO。

  • Epistemic Mode: ASSERTED
  • Mapping Status: PARTIAL

③ Semantic Frame

{
  "attribute": "INTERACTIVITY",
  "context_nodes": [
    {
      "id": "04_knowledge_base/NVIDIA B200",
      "label": "B200"
    }
  ],
  "entity": {
    "id": "04_knowledge_base/Multi-Token Prediction",
    "label": "MTP"
  },
  "frame_type": "ATTRIBUTE",
  "qualifiers": {
    "condition_text": null,
    "numeric_mentions": [
      "8k"
    ],
    "temporal_mentions": []
  },
  "value": {
    "numeric_mentions": [
      "8k"
    ],
    "value_text": "在此 interactivity 水準(8k1k workload)下,開啟 MTP 的 B200 可以達到最佳 perf/TCO。"
  }
}

④ Canonical Entity Mapping

RoleSurface LabelCanonical Target
entityMTP04_knowledge_base/Multi-Token Prediction
context_0B20004_knowledge_base/NVIDIA B200

⑤ Human Review

請在 Properties 逐項確認:

  • 原文 → Atomic Claim 是否忠實
  • Atomic Claim → Semantic Frame 是否忠實
  • Canonical Entity mapping 是否正確
  • Epistemic mode 是否保留原文語氣
  • 最後選擇 review_action

Review state

Markdown 內文不是正式 approval。只有 Apply bridge 寫入的 Decision Ledger event 才是正式決策。