IX2-0305

① SA Source

Context Before

AMD has launched a new inference engine called ATOM. Atom can deliver slightly better single node performance, but it is completely lacking on a lot of features that makes it unusable for real workloads. One such example is that it does not support NVMe or CPU KVCache offloading, tool parsing, wide expert parallelism, or disaggregated serving. This has led to zero customers using it in production. Unlike Nvidia’s TRTLLM which generates billions of tokens per hour globally at companies like TogetherAI, etc and does support tool parsing and other features , there are no token factories currently using ATOM due to the lack of the aforementioned features.

Furthermore, maintainers of open-source inference engines like vLLM are disappointed in AMD due to a lack of engineering and GPU resources provided by AMD. For example, Simon Mo, lead vLLM maintainer, states in this GitHub RFC that there is still no working MI355X that he can add to vLLM CI, hence the poor user experience. There are currently zero Mi355X tests on vLLM, while NVIDIA’s B200 has many tests on vLLM. Similarly, there are still not enough MI300X CI machines on vLLM. Upstream vLLM needs at least 20 more MI300 machines, 20 more MI325 machines and 20 more MI355X machines to reach the same level of usability as CUDA.

Evidence

We will talk more about AMD’s previous lackluster contribution towards vLLM, SGLang, PyTorch CI machine situation & how Anush started to fix it in our upcoming State of AMD article

Context After

Moreover, the vLLM maintainers say that they cannot support day 0 vLLM support for ROCm due to this issue of lack of machine resources. This huge disparity in time to market continues to lead to ROCm lagging behind and leaving a huge opening for Nvidia to continue to charge an insane 75% gross margin (4x markup on cost of goods).

image

② Atomic Claim

SemiAnalysis 將在後續「State of AMD」文章中,更深入討論 AMD 過去對 vLLMSGLangPyTorch CI machines 貢獻不足,以及 Anush 如何開始改善這個問題。

  • Epistemic Mode: ASSERTED
  • Mapping Status: COMPLETE

③ Semantic Frame

{
  "attribute": "COUNT",
  "context_nodes": [
    {
      "id": "04_knowledge_base/vLLM",
      "label": "vLLM"
    },
    {
      "id": "04_knowledge_base/SGLang",
      "label": "SGLang"
    },
    {
      "id": "04_knowledge_base/PyTorch",
      "label": "PyTorch"
    }
  ],
  "entity": {
    "id": "02_companies/AMD",
    "label": "AMD"
  },
  "frame_type": "ATTRIBUTE",
  "qualifiers": {
    "condition_text": null,
    "numeric_mentions": [],
    "temporal_mentions": []
  },
  "value": {
    "numeric_mentions": [],
    "value_text": "SemiAnalysis 將在後續「State of AMD」文章中,更深入討論 AMD 過去對 vLLM、SGLang、PyTorch CI machines 貢獻不足,以及 Anush 如何開始改善這個問題。"
  }
}

④ Canonical Entity Mapping

RoleSurface LabelCanonical Target
entityAMDAMD
context_0vLLMvLLM
context_1SGLangSGLang
context_2PyTorchPyTorch

⑤ Human Review

請在 Properties 逐項確認:

  • 原文 → Atomic Claim 是否忠實
  • Atomic Claim → Semantic Frame 是否忠實
  • Canonical Entity mapping 是否正確
  • Epistemic mode 是否保留原文語氣
  • 最後選擇 review_action

Review state

Markdown 內文不是正式 approval。只有 Apply bridge 寫入的 Decision Ledger event 才是正式決策。