IX2-0020

① SA Source

Context Before

Source: InferenceMAX GitHub

Key Observations and Results to Highlight

Evidence

when compared to widely used Dynamo TRTLLM B200 FP8, TRT continues to framemog

Context After

We also see that for single node aggregated serving, AMD’s SGLang delivers better perf per TCO than NVIDIA’s SGLang for FP8. It is also great to see that AMD has deprecated their second class fork of vllm to move further upstream and closer to delivering first class experience. Stay tuned for our “State of AMD” article where we talk about the many areas where AMD’s pace of improvement has been rapid & also the areas where the pace of improvement has been lackluster. We recommend that NVIDIA focus even more on SGLang & vLLM ecosystem in addition their TRTLLM engine. Jensen needs to staff more resources & engineers towards contributing open ecosystems like SGLang & vLLM .

SemiAnalysis InferenceX is free open source software and reader-supported. To receive new posts and support our work consider becoming a free or paid subscriber.

② Atomic Claim

但若拿來與廣泛使用的 Dynamo TRTLLM B200 FP8 比較,TRT 仍明顯領先。

  • Epistemic Mode: ASSERTED
  • Mapping Status: COMPLETE

③ Semantic Frame

{
  "frame_type": "NARY_RELATION",
  "participants": [
    {
      "node": {
        "id": "04_knowledge_base/TensorRT-LLM",
        "label": "TRTLLM"
      },
      "role": "user_or_subject"
    },
    {
      "node": {
        "id": "04_knowledge_base/NVIDIA B200",
        "label": "B200"
      },
      "role": "used_entity"
    },
    {
      "node": {
        "id": "04_knowledge_base/FP8",
        "label": "FP8"
      },
      "role": "used_entity"
    }
  ],
  "qualifiers": {
    "condition_text": null,
    "numeric_mentions": [],
    "temporal_mentions": []
  },
  "relation_type": "USES"
}

④ Canonical Entity Mapping

RoleSurface LabelCanonical Target
user_or_subjectTRTLLMTensorRT-LLM
used_entityB20004_knowledge_base/NVIDIA B200
used_entityFP8FP8

⑤ Human Review

請在 Properties 逐項確認:

  • 原文 → Atomic Claim 是否忠實
  • Atomic Claim → Semantic Frame 是否忠實
  • Canonical Entity mapping 是否正確
  • Epistemic mode 是否保留原文語氣
  • 最後選擇 review_action

Review state

Markdown 內文不是正式 approval。只有 Apply bridge 寫入的 Decision Ledger event 才是正式決策。