IX2-0189

① SA Source

Context Before

image

Source: SemiAnalysis InferenceX

Evidence

Let’s use the findings above to dig deeper into the unit economics of serving LLMs at scale. From the OpenRouter data above, we see that Crusoe serves at 36 tok/sec/user at 5.40/M output tokens. If we assume no cache hits and that Crusoe is using at least H200s with SOTA inference techniques like MTP, disagg, and wide EP, the data above suggests they incur a cost of _no more than _/M input tokens and $2.955/M output tokens for a profit margin of up to 83% gross margin (depreciation counted in cost of goods sold) on input tokens and 45% gross margin on output tokens.

Context After

SemiAnalysis InferenceX is free open source software and reader-supported. To receive new posts and support our work, consider becoming a free or paid subscriber.

Subscribed

② Atomic Claim

假設 Crusoe 無 cache hit、至少使用 H200,並採 MTP、disaggregated serving、wide EP 等 SOTA inference techniques,SemiAnalysis 估計 input token cost 不超過 2.955/M;對應 gross margin 分別最高約 83% 與 45%,折舊計入 COGS。

  • Epistemic Mode: HYPOTHETICAL
  • Mapping Status: PARTIAL

③ Semantic Frame

{
  "attribute": "SERVING_UNIT_ECONOMICS",
  "context_nodes": [
    {
      "id": "04_knowledge_base/H200",
      "label": "H200"
    },
    {
      "id": "04_knowledge_base/Multi-Token Prediction",
      "label": "MTP"
    },
    {
      "id": "04_knowledge_base/Expert Parallelism",
      "label": "wide EP"
    }
  ],
  "entity": {
    "id": "02_companies/Crusoe",
    "label": "Crusoe"
  },
  "frame_type": "ATTRIBUTE",
  "qualifiers": {
    "condition_text": "no cache hits; at least H200; SOTA inference techniques including MTP, disagg, wide EP",
    "numeric_mentions": [
      "$0.226/M",
      "$2.955/M",
      "83%",
      "45%"
    ],
    "temporal_mentions": []
  },
  "value": {
    "numeric_mentions": [
      "$0.226/M",
      "$2.955/M",
      "83%",
      "45%"
    ],
    "value_text": "假設 [[02_companies/Crusoe|Crusoe]] 無 cache hit、至少使用 [[04_knowledge_base/H200|H200]],並採 [[04_knowledge_base/Multi-Token Prediction|MTP]]、disaggregated serving、[[04_knowledge_base/Expert Parallelism|wide EP]] 等 SOTA inference techniques,SemiAnalysis 估計 input token cost 不超過 $0.226/M、output token cost 不超過 $2.955/M;對應 gross margin 分別最高約 83% 與 45%,折舊計入 COGS。"
  }
}

④ Canonical Entity Mapping

RoleSurface LabelCanonical Target
entityCrusoeCrusoe
context_0H200H200
context_1MTP04_knowledge_base/Multi-Token Prediction
context_2wide EP04_knowledge_base/Expert Parallelism

⑤ Human Review

請在 Properties 逐項確認:

  • 原文 → Atomic Claim 是否忠實
  • Atomic Claim → Semantic Frame 是否忠實
  • Canonical Entity mapping 是否正確
  • Epistemic mode 是否保留原文語氣
  • 最後選擇 review_action

Review state

Markdown 內文不是正式 approval。只有 Apply bridge 寫入的 Decision Ledger event 才是正式決策。