IX2-0176

① SA Source

Context Before

image

Source: OpenRouter

Evidence

We can then use real InferenceX data to interpolate the cost per million input/output tokens at an interactivity level of 35 tok/sec/user, which is a reasonable interactivity level given the data above

Context After

As we mention later in the article, this is best understood as _baseline _data and not completely representative of real-world inference, mainly because InferenceX benchmarks on random data and disables prefix caching. In other words, performance/cost will be _at least _this good. It is also important to note that there are not data points for each GPU at _each _interactivity level. Thus we cannot make _exact _comparisons at each degree of interactivity. We nevertheless think the bar chart comparisons presented below are (very) reasonable interpolations in lieu of using exact data points.

Comparing disagg+wideEP configs at this interactivity level, we see just how effective distributed inference techniques are when it comes to both perf/TCO and overall throughput. We also see how large scale up domains (like GB300 and GB200 NVL72) absolutely dominate in total throughput per GPU.

② Atomic Claim

因此可以利用真實 InferenceX 資料,插值估算在 35 tok/sec/user interactivity 下每百萬 input/output tokens 的成本;根據上述資料,這是一個合理 interactivity 水準。

  • Epistemic Mode: ASSERTED
  • Mapping Status: PARTIAL

③ Semantic Frame

{
  "attribute": "THROUGHPUT",
  "context_nodes": [],
  "entity": {
    "id": "04_knowledge_base/InferenceX",
    "label": "InferenceX"
  },
  "frame_type": "ATTRIBUTE",
  "qualifiers": {
    "condition_text": null,
    "numeric_mentions": [
      "35 tok/s"
    ],
    "temporal_mentions": []
  },
  "value": {
    "numeric_mentions": [
      "35 tok/s"
    ],
    "value_text": "因此可以利用真實 InferenceX 資料,插值估算在 35 tok/sec/user interactivity 下每百萬 input/output tokens 的成本;根據上述資料,這是一個合理 interactivity 水準。"
  }
}

④ Canonical Entity Mapping

RoleSurface LabelCanonical Target
entityInferenceXInferenceX

⑤ Human Review

請在 Properties 逐項確認:

  • 原文 → Atomic Claim 是否忠實
  • Atomic Claim → Semantic Frame 是否忠實
  • Canonical Entity mapping 是否正確
  • Epistemic mode 是否保留原文語氣
  • 最後選擇 review_action

Review state

Markdown 內文不是正式 approval。只有 Apply bridge 寫入的 Decision Ledger event 才是正式決策。