IX2-0176
① SA Source
- Source: 開啟完整 SA 文章
- Section:
Unpacking Inference Providers’ Unit Economics - Line hint:
312
Context Before

Source: OpenRouter ↗
Evidence
We can then use real InferenceX data to interpolate the cost per million input/output tokens at an interactivity level of 35 tok/sec/user, which is a reasonable interactivity level given the data above
Context After
As we mention later in the article, this is best understood as _baseline _data and not completely representative of real-world inference, mainly because InferenceX benchmarks on random data and disables prefix caching. In other words, performance/cost will be _at least _this good. It is also important to note that there are not data points for each GPU at _each _interactivity level. Thus we cannot make _exact _comparisons at each degree of interactivity. We nevertheless think the bar chart comparisons presented below are (very) reasonable interpolations in lieu of using exact data points.
Comparing disagg+wideEP configs at this interactivity level, we see just how effective distributed inference techniques are when it comes to both perf/TCO and overall throughput. We also see how large scale up domains (like GB300 and GB200 NVL72) absolutely dominate in total throughput per GPU.
② Atomic Claim
因此可以利用真實 InferenceX 資料,插值估算在 35 tok/sec/user interactivity 下每百萬 input/output tokens 的成本;根據上述資料,這是一個合理 interactivity 水準。
- Epistemic Mode:
ASSERTED - Mapping Status:
PARTIAL
③ Semantic Frame
{
"attribute": "THROUGHPUT",
"context_nodes": [],
"entity": {
"id": "04_knowledge_base/InferenceX",
"label": "InferenceX"
},
"frame_type": "ATTRIBUTE",
"qualifiers": {
"condition_text": null,
"numeric_mentions": [
"35 tok/s"
],
"temporal_mentions": []
},
"value": {
"numeric_mentions": [
"35 tok/s"
],
"value_text": "因此可以利用真實 InferenceX 資料,插值估算在 35 tok/sec/user interactivity 下每百萬 input/output tokens 的成本;根據上述資料,這是一個合理 interactivity 水準。"
}
}④ Canonical Entity Mapping
| Role | Surface Label | Canonical Target |
|---|---|---|
| entity | InferenceX | InferenceX |
⑤ Human Review
請在 Properties 逐項確認:
- 原文 → Atomic Claim 是否忠實
- Atomic Claim → Semantic Frame 是否忠實
- Canonical Entity mapping 是否正確
- Epistemic mode 是否保留原文語氣
- 最後選擇
review_action
Review state
Markdown 內文不是正式 approval。只有 Apply bridge 寫入的 Decision Ledger event 才是正式決策。