IX2-0003

① SA Source

Context Before

Introduction

InferenceXv2 (formerly InferenceMAX) builds on the foundation established by InferenceMAXv1, our open-source, continuously updated inference benchmark that has set a new standard for AI inference performance and economics. InferenceMAXv1 moved beyond static, point-in-time benchmarks by running continuous tests across hundreds of chips and popular open-source frameworks. Free dashboard available here.

Evidence

Our benchmark has been widely reproduced, validated and/or supported by almost every major buyer of compute from Google Cloud to Microsoft Azure to Oracle, OpenAI , and many more.

Context After

InferenceXv2 builds on this foundation. It expands coverage to include large scale DeepSeek MoE disaggregated inference (disagg prefill, or simply “disagg”) with wide expert parallelism (wideEP) optimization to **all 6 NVIDIA western GPU SKUs from the past 4 years **as well as to every single AMD western GPU SKU released in the past 3 years – in total InferenceXv2 utilizes close to 1000 frontier GPUs for a full benchmark run across all SKUs.

With today’s release, InferenceXv2 is now the first suite to benchmark the Blackwell Ultra GB300 NVL72 and B300 across the whole pareto frontier curve, and it is the first third party benchmark to test disagg+wideEP multi-node FP4 and FP8 MI355X performance. In future iterations of InferenceX, we will continue to focus heavily on disaggregated serving with wide expert parallelism as that is what is deployed in production at Frontier AI Labs like OpenAI, Anthropic, xAI, Google Deepmind, DeepSeek as well as advanced API providers like TogetherAI, Baseten, and Fireworks. In this article, we will also break down the system engineering principles and economics in play around the latest Claude Code Fast mode feature .

② Atomic Claim

這套 benchmark 已被幾乎所有主要算力買方廣泛重現、驗證或支持,包括 Google Cloud、Microsoft AzureOracleOpenAI 等。

  • Epistemic Mode: ASSERTED
  • Mapping Status: PARTIAL

③ Semantic Frame

{
  "attribute": "UNSPECIFIED_ATTRIBUTE",
  "context_nodes": [
    {
      "id": "02_companies/MSFT",
      "label": "Microsoft"
    },
    {
      "id": "04_knowledge_base/Microsoft Azure",
      "label": "Azure"
    },
    {
      "id": "02_companies/ORCL",
      "label": "Oracle"
    },
    {
      "id": "02_companies/OpenAI",
      "label": "OpenAI"
    }
  ],
  "entity": {
    "id": "02_companies/GOOG",
    "label": "Google"
  },
  "frame_type": "ATTRIBUTE",
  "qualifiers": {
    "condition_text": null,
    "numeric_mentions": [],
    "temporal_mentions": []
  },
  "value": {
    "numeric_mentions": [],
    "value_text": "這套 benchmark 已被幾乎所有主要算力買方廣泛重現、驗證或支持,包括 Google Cloud、Microsoft Azure、Oracle、OpenAI 等。"
  }
}

④ Canonical Entity Mapping

RoleSurface LabelCanonical Target
entityGoogleGOOG
context_0MicrosoftMSFT
context_1Azure04_knowledge_base/Microsoft Azure
context_2OracleORCL
context_3OpenAIOpenAI

⑤ Human Review

請在 Properties 逐項確認:

  • 原文 → Atomic Claim 是否忠實
  • Atomic Claim → Semantic Frame 是否忠實
  • Canonical Entity mapping 是否正確
  • Epistemic mode 是否保留原文語氣
  • 最後選擇 review_action

Review state

Markdown 內文不是正式 approval。只有 Apply bridge 寫入的 Decision Ledger event 才是正式決策。