IX2-0302
① SA Source
- Source: 開啟完整 SA 文章
- Section:
AMD ATOM Engine - Line hint:
496
Context Before
AMD ATOM Engine
AMD has launched a new inference engine called ATOM. Atom can deliver slightly better single node performance, but it is completely lacking on a lot of features that makes it unusable for real workloads. One such example is that it does not support NVMe or CPU KVCache offloading, tool parsing, wide expert parallelism, or disaggregated serving. This has led to zero customers using it in production. Unlike Nvidia’s TRTLLM which generates billions of tokens per hour globally at companies like TogetherAI, etc and does support tool parsing and other features ↗, there are no token factories currently using ATOM due to the lack of the aforementioned features.
Evidence
Context After
We at SemiAnalysis have been trying to get AMD to contribute more compute to vLLM and have had some success on that within the couple weeks. vLLM will start to get a couple of MI355X machines such that they can bring their CI test parity from 0% to non-0%. We will talk more about AMD’s previous lackluster contribution towards vLLM, SGLang, PyTorch CI machine situation & how Anush started to fix it in our upcoming State of AMD article. At SemiAnalysis, we will have internal dashboard to track the # of tests & quality of tests that AMD & NVIDIA runs on vLLM, SGLang, PyTorch, & JAX.
Moreover, the vLLM maintainers say that they cannot support day 0 vLLM support for ROCm due to this issue of lack of machine resources. This huge disparity in time to market continues to lead to ROCm lagging behind and leaving a huge opening for Nvidia to continue to charge an insane 75% gross margin (4x markup on cost of goods).
② Atomic Claim
Upstream vLLM 至少還需要 20 台 MI300、20 台 MI325 與 20 台 MI355X machines,才能達到接近 CUDA 的可用性水準。
- Epistemic Mode:
ASSERTED - Mapping Status:
COMPLETE
③ Semantic Frame
{
"attribute": "LIFECYCLE_STATUS",
"context_nodes": [
{
"id": "04_knowledge_base/MI300",
"label": "MI300"
},
{
"id": "04_knowledge_base/MI325",
"label": "MI325"
},
{
"id": "04_knowledge_base/MI355X",
"label": "MI355X"
},
{
"id": "04_knowledge_base/CUDA",
"label": "CUDA"
}
],
"entity": {
"id": "04_knowledge_base/vLLM",
"label": "vLLM"
},
"frame_type": "ATTRIBUTE",
"qualifiers": {
"condition_text": null,
"numeric_mentions": [
"20"
],
"temporal_mentions": []
},
"value": {
"numeric_mentions": [
"20"
],
"value_text": "Upstream vLLM 至少還需要 20 台 MI300、20 台 MI325 與 20 台 MI355X machines,才能達到接近 CUDA 的可用性水準。"
}
}④ Canonical Entity Mapping
| Role | Surface Label | Canonical Target |
|---|---|---|
| entity | vLLM | vLLM |
| context_0 | MI300 | MI300 |
| context_1 | MI325 | MI325 |
| context_2 | MI355X | MI355X |
| context_3 | CUDA | CUDA |
⑤ Human Review
請在 Properties 逐項確認:
- 原文 → Atomic Claim 是否忠實
- Atomic Claim → Semantic Frame 是否忠實
- Canonical Entity mapping 是否正確
- Epistemic mode 是否保留原文語氣
- 最後選擇
review_action
Review state
Markdown 內文不是正式 approval。只有 Apply bridge 寫入的 Decision Ledger event 才是正式決策。