2025-08-12_scaling-the-memory-wall-the-rise-and-roadmap-of-hbm::HBM25-0025

① SA Source

Context Before

Bandwidth beats capacity

Although every accelerator has strived to design in the most HBM that is attainable, we understand that OpenAI’s ASIC project is going to break this trend. OpenAI is choosing to use 8-Hi HBM4 instead of opting for 16-Hi or even 12-Hi HBM. 8-Hi has been expected to be phased out as it is not mentioned on the major memory vendors’ roadmaps. 12-Hi becoming the standard from as early as next year. This is notable as the first time a customer that is aggressive on silicon has requested a downgrade in specifications.

Evidence

This is because OAI sees 8-Hi offering a much better ratio of bandwidth to capacity at the given cost. Capacity is important, but for inference bandwidth is most often the constraint. With 8-Hi stacks, OAI gets the same bandwidth but at less than half the price per stack.

Context After

However, this shouldn’t be interpreted as a leading lab calling time on the trend of HBM capacity scaling. OAI will focus more on improving the software, micro-architecture, networking within the broader system so that future generations can become performace/TCO competitive with merchant solutions. Reducing memory costs is a good way to reduce the investment without compromising on the other surface areas for improvement and adding more capacity with higher layers is relatively trivial. OpenAI will still be dependent on GPUs in that timeframe. Rubin Ultra will be there with abundant capacity as model architects will no doubt try to find ways to consume this additional capacity to extract more intelligence. Going for lower capacity is a cheaper, and relative lower risk way for OpenAI to explore different cost and performance tradeoffs for their accelerator.

Ultimately HBM demand will be influenced by both user behavior and the architectural decisions of accelerator designers, ie what they consider is optimal in terms of HBM memory capacity / bandwidth vs FLOPs over the lifetime of the chip. With a lifecycle of 4 years, designers need to take into account the evolution of workloads over that period and make design choices that are flexible enough to adapt to different needs of model architectures for inference and training.

② Atomic Claim

OpenAI 認為 8-Hi HBM4 在 inference 中可維持相同 bandwidth、但每 stack 成本不到一半,因此 bandwidth/capacity/cost ratio 更佳。

  • Epistemic Mode: ATTRIBUTED
  • Mapping Status: PARTIAL

③ Semantic Frame

{
  "attribute": "8_hi_cost_bandwidth_tradeoff",
  "context_nodes": [
    {
      "id": "02_companies/OpenAI",
      "label": "OpenAI"
    }
  ],
  "entity": {
    "id": "04_knowledge_base/HBM4",
    "label": "HBM4"
  },
  "frame_type": "ATTRIBUTE",
  "qualifiers": {
    "condition_text": "inference bandwidth is often the constraint",
    "numeric_mentions": [
      "8-Hi"
    ],
    "temporal_mentions": []
  },
  "value": {
    "numeric_mentions": [
      "8-Hi"
    ],
    "value_text": "same bandwidth at less than half the price per stack"
  }
}

④ Canonical Entity Mapping

RoleSurface LabelCanonical Target
entityHBM4HBM4
context_0OpenAIOpenAI

⑤ Human Review

請在 Properties 逐項確認:

  • 原文 → Atomic Claim 是否忠實
  • Atomic Claim → Semantic Frame 是否忠實
  • Canonical Entity mapping 是否正確
  • Epistemic mode 是否保留原文語氣
  • 最後選擇 review_action

Review state

Markdown 內文不是正式 approval。只有 Apply bridge 寫入的 Decision Ledger event 才是正式決策。