IX2-0596

① SA Source

Context Before

With the rise of Claude Code, Codex, and Kimi, it is becoming increasingly important to benchmark performance in agentic coding scenarios. Like above, these scenarios are multi-turn but also include extremely long context conversations as well as tool use. In the next few months, we plan on creating a benchmark suite that will most accurately capture the performance of open models in these agentic coding scenarios across all chips.

Adding TPU, Trainium and More Models

Evidence

Currently, we continuously benchmark DeepSeek R1 and GPT OSS 120B (previously Llama 3.1 70B as well)

Context After

In addition to new models, we are actively working on adding both TPU and Trainium.

Total Cost of Ownership (NVL72, Blackwell, Blackwell Ultra, MI355, Hopper, MI325, MI300)

② Atomic Claim

目前團隊持續 benchmark DeepSeek R1 與 GPT OSS 120B;過去也包含 Llama 3.1 70B。

  • Epistemic Mode: ASSERTED
  • Mapping Status: COMPLETE

③ Semantic Frame

{
  "additional_nodes": [],
  "frame_type": "RELATION",
  "object": {
    "id": "04_knowledge_base/Llama",
    "label": "Llama 3"
  },
  "predicate": "HAS_COMPONENT",
  "qualifiers": {
    "condition_text": null,
    "numeric_mentions": [
      "120B",
      "3.1",
      "70B"
    ],
    "temporal_mentions": []
  },
  "subject": {
    "id": "02_companies/DeepSeek",
    "label": "DeepSeek"
  }
}

④ Canonical Entity Mapping

RoleSurface LabelCanonical Target
subjectDeepSeekDeepSeek
objectLlama 3Llama

⑤ Human Review

請在 Properties 逐項確認:

  • 原文 → Atomic Claim 是否忠實
  • Atomic Claim → Semantic Frame 是否忠實
  • Canonical Entity mapping 是否正確
  • Epistemic mode 是否保留原文語氣
  • 最後選擇 review_action

Review state

Markdown 內文不是正式 approval。只有 Apply bridge 寫入的 Decision Ledger event 才是正式決策。