IX2-0011::SPLIT02
① SA Source
- Source: 開啟完整 SA 文章
- Section:
Introduction - Line hint:
25
Context Before
Our benchmark has been widely reproduced, validated and/or supported by almost every major buyer ↗ of compute from Google Cloud ↗ to Microsoft Azure ↗ to Oracle, OpenAI ↗, and many more.
InferenceXv2 builds on this foundation. It expands coverage to include large scale DeepSeek MoE disaggregated inference (disagg prefill, or simply “disagg”) with wide expert parallelism (wideEP) optimization to **all 6 NVIDIA western GPU SKUs from the past 4 years **as well as to every single AMD western GPU SKU released in the past 3 years – in total InferenceXv2 utilizes close to 1000 frontier GPUs for a full benchmark run across all SKUs.
Evidence
disaggregated serving with wide expert parallelism as that is what is deployed in production at Frontier AI Labs like OpenAI, Anthropic, xAI, Google Deepmind, DeepSeek as well as advanced API providers like TogetherAI, Baseten, and Fireworks
Context After
Our benchmark is completely open-source under Apache 2.0 – this means that we are able to move at the same rapid speed at which the AI software ecosystem is advancing. If you like our work and would like to show us some support, please drop a star on our GitHub ↗! We also provide a free data visualizer at https://inferencex.com ↗ for everyone in the ML community to explore the complete dataset themselves.
We will add DeepSeekv4 and other popular Chinese frontier models with day 0 support as over the past 6 months, we now have cleaned up a lot of tech debt and are able to move fast with stable infrastructure ↗. We will also be adding TPUv7 Ironwood and Trainium3 to InferenceX later this year! If you want to contribute to our impactful mission while earning a competitive compensation, consider applying here ↗.
② Atomic Claim
OpenAI、Anthropic、xAI、Google Deepmind、DeepSeek 等 Frontier AI Labs,以及 TogetherAI、Baseten、Fireworks 等 API providers,在 production 中部署搭配 wide expert parallelism 的 disaggregated serving。
- Epistemic Mode:
ASSERTED - Mapping Status:
PARTIAL
③ Semantic Frame
{
"frame_type": "NARY_RELATION",
"participants": [
{
"node": {
"id": "04_knowledge_base/Disaggregated serving",
"label": "disaggregated serving"
},
"role": "technique"
},
{
"node": {
"id": "04_knowledge_base/Expert Parallelism",
"label": "wide expert parallelism"
},
"role": "technique"
},
{
"node": {
"id": "02_companies/OpenAI",
"label": "OpenAI"
},
"role": "deployer"
},
{
"node": {
"id": "02_companies/Anthropic",
"label": "Anthropic"
},
"role": "deployer"
},
{
"node": {
"id": "02_companies/xAI",
"label": "xAI"
},
"role": "deployer"
},
{
"node": {
"id": "02_companies/GOOG",
"label": "Google"
},
"role": "deployer"
},
{
"node": {
"id": "02_companies/DeepMind",
"label": "Deepmind"
},
"role": "deployer"
},
{
"node": {
"id": "02_companies/DeepSeek",
"label": "DeepSeek"
},
"role": "deployer"
},
{
"node": {
"id": "02_companies/Baseten",
"label": "Baseten"
},
"role": "deployer"
}
],
"qualifiers": {
"condition_text": "production deployment; TogetherAI and Fireworks remain unmapped text participants",
"numeric_mentions": [],
"temporal_mentions": []
},
"relation_type": "DEPLOYS"
}④ Canonical Entity Mapping
| Role | Surface Label | Canonical Target |
|---|---|---|
| technique | disaggregated serving | 04_knowledge_base/Disaggregated serving |
| technique | wide expert parallelism | 04_knowledge_base/Expert Parallelism |
| deployer | OpenAI | OpenAI |
| deployer | Anthropic | Anthropic |
| deployer | xAI | xAI |
| deployer | GOOG | |
| deployer | Deepmind | DeepMind |
| deployer | DeepSeek | DeepSeek |
| deployer | Baseten | Baseten |
⑤ Human Review
請在 Properties 逐項確認:
- 原文 → Atomic Claim 是否忠實
- Atomic Claim → Semantic Frame 是否忠實
- Canonical Entity mapping 是否正確
- Epistemic mode 是否保留原文語氣
- 最後選擇
review_action
Review state
Markdown 內文不是正式 approval。只有 Apply bridge 寫入的 Decision Ledger event 才是正式決策。