IX2-0183
① SA Source
- Source: 開啟完整 SA 文章
- Section:
Unpacking Inference Providers’ Unit Economics - Line hint:
316
Context Before
We can then use real InferenceX data to interpolate the cost per million input/output tokens at an interactivity level of 35 tok/sec/user, which is a reasonable interactivity level given the data above.
As we mention later in the article, this is best understood as _baseline _data and not completely representative of real-world inference, mainly because InferenceX benchmarks on random data and disables prefix caching. In other words, performance/cost will be _at least _this good. It is also important to note that there are not data points for each GPU at _each _interactivity level. Thus we cannot make _exact _comparisons at each degree of interactivity. We nevertheless think the bar chart comparisons presented below are (very) reasonable interpolations in lieu of using exact data points.
Evidence
We also see how large scale up domains (like GB300 and GB200 NVL72) absolutely dominate in total throughput per GPU
Context After
It is interesting to note that at this interactivity level (on an 8k1k workload type), the B200 can achieve the best perf/TCO when MTP is enabled. Below we also list the Total Cost of Ownership (TCO) (Owning – Hyperscaler) for each GPU:

② Atomic Claim
也可以看到像 GB300 與 GB200 NVL72 這類大型 scale-up domains,在每 GPU 總 throughput 上具有明顯優勢。
- Epistemic Mode:
ASSERTED - Mapping Status:
COMPLETE
③ Semantic Frame
{
"attribute": "THROUGHPUT",
"context_nodes": [
{
"id": "04_knowledge_base/GB200 NVL72",
"label": "GB200 NVL72"
},
{
"id": "04_knowledge_base/GPU",
"label": "GPU"
}
],
"entity": {
"id": "04_knowledge_base/GB300",
"label": "GB300"
},
"frame_type": "ATTRIBUTE",
"qualifiers": {
"condition_text": null,
"numeric_mentions": [],
"temporal_mentions": []
},
"value": {
"numeric_mentions": [],
"value_text": "也可以看到像 GB300 與 GB200 NVL72 這類大型 scale-up domains,在每 GPU 總 throughput 上具有明顯優勢。"
}
}④ Canonical Entity Mapping
| Role | Surface Label | Canonical Target |
|---|---|---|
| entity | GB300 | GB300 |
| context_0 | GB200 NVL72 | 04_knowledge_base/GB200 NVL72 |
| context_1 | GPU | GPU |
⑤ Human Review
請在 Properties 逐項確認:
- 原文 → Atomic Claim 是否忠實
- Atomic Claim → Semantic Frame 是否忠實
- Canonical Entity mapping 是否正確
- Epistemic mode 是否保留原文語氣
- 最後選擇
review_action
Review state
Markdown 內文不是正式 approval。只有 Apply bridge 寫入的 Decision Ledger event 才是正式決策。