已核验 · Aug 5, 2026
已独立佐证推理硬件集群:Blackwell + Groq LPU v3 + Cerebras WSE + Apple MLX 1.0 + Inferentia3 + TPU v6 —— 每个 vendor 都发货 Blackwell-era 答案
8 个信源2026-08-04 这一周出现了一组推理硬件 / 框架发布集群:NVIDIA Blackwell B200(FP4/FP8 + NVLink Switch + DGX SuperPOD)、Groq LPU 推理引擎 v3(无论 batch 大小都确定性 token 流延迟)、Cerebras WSE(晶圆级单芯片推理 + cloud SDK)、Apple MLX 1.0 + MLX-LM(Apple Silicon 设备端 LLM 推理,统一内存)、AWS Inferentia3 + Neuron SDK v3(Neuron-compatible 实例生产推理)、Google TPU v6 Trillium(Vertex AI 托管推理)。The Decoder 把这个集群框成「每个推理 vendor 都发货一个 Blackwell-era 答案」—— GPU 代际成了竞争定位单位,每个 vendor 都挑了一条路。
为什么现在讲
集群是把六个独立推理发布变成一个连贯「推理现在是一个 buyer-side 市场」故事的 editorial framing —— 并排对比对创作者最有用。
推荐理由
演示空间:同一推理任务(token 流延迟、吞吐、成本)在 Blackwell、Groq LPU、Cerebras WSE、Apple MLX、Inferentia3、TPU v6 上的并排对比。
依据
The Decoder + IT 之家周报 + 六个独立 vendor 一手源
“六个推理 vendor 同一周发了 Blackwell-era 答案 —— NVIDIA Blackwell、Groq LPU v3、Cerebras WSE、Apple MLX 1.0、AWS Inferentia3、Google TPU v6 —— 每个 vendor 在 cost、latency、model coverage、设备端支持上都挑了不同的路。”
切入角度
用集群引入「推理现在是一个 buyer-side 市场」模式 —— 每个 vendor 在 cost / latency / model coverage / 设备端上都挑了一条路 —— 用这个 lens 并排对比 vendor 方案。
形式
长视频讲解
演示想法
录一段 16 分钟并排对比讲解:2 分钟开场讲「推理 buyer-side 市场」框架,然后每个 vendor 2 分钟(Blackwell / Groq / Cerebras / Apple MLX / Inferentia3 / TPU v6),最后 4 分钟在同一 token 流延迟 / 吞吐 / 成本测量上并排对比六个。
平台注意
每个 vendor 都把发布框成跟自己关心的竞争集对比;The Decoder 和 IT 之家是 editorial framing 层,不是独立验证。在记录任何具体吞吐或定价声明前,对照 vendor 一手文档确认。
可用说法
- NVIDIA Blackwell generation (B200 datasheet) documents FP4/FP8 inference throughput, NVLink Switch fabric scaling, and DGX SuperPOD reference architecture for production inference.
- Groq LPU inference engine v3 positions deterministic token-stream latency as the headline product property, regardless of batch size.
- Cerebras WSE wafer-scale inference ships with cloud SDK for inference, avoiding model parallelism across GPUs via single-chip inference.
- Apple MLX 1.0 framework GA + MLX-LM Python package for on-device LLM inference on Apple Silicon, with unified memory sharing GPU and CPU memory.
- AWS Inferentia3 GA + Neuron SDK v3 with PyTorch / JAX / TensorFlow support, distributed inference, and the Neuron Compiler for custom model compilation.
- Google Cloud TPU v6 (Trillium) GA on Cloud with Vertex AI managed inference integration and per-pod scaling topology.
证据链
来自新闻
拆解
六个推理 vendor 同一周发了 Blackwell-era 答案 —— editorial framing(「每个推理 vendor 都发货 Blackwell-era 答案」)有用,但如果你不引入 buyer-side 模式,内容就退化成 spec sheet。本篇解释怎么用集群引入「推理现在是一个 buyer-side 市场」模式(每个 vendor 在 cost / latency / model coverage / 设备端上都挑了一条路),用这个 lens 并排对比 vendor 方案。
信源
- The Decoder: tech-press coverage of the 'inference hardware' cluster for the week of 2026-08-04
- IT之家: 中文科技媒体覆盖 8/4 推理硬件集群
- NVIDIA: Blackwell B200 GPU datasheet and inference benchmarks
- Groq: LPU inference engine v3 — deterministic token-stream latency
- Cerebras: WSE wafer-scale inference + cloud SDK
- Apple: MLX 1.0 framework + MLX-LM for on-device inference
- AWS: Inferentia3 + Neuron SDK v3 GA for production inference
- Google Cloud: TPU v6 (Trillium) GA on Cloud + Vertex AI integration
风险
- Use The Decoder and IT之家 as media-type corroboration, but read the underlying vendor docs for any specific throughput or pricing claim before stating it on the record. Verify specific capability claim against the underlying vendor docs and the actual license / pricing matrix before stating it on the record; do not paraphrase per-platform pricing or license terms into specific dollar figures or commercial-use clauses.
- Vendor docs confirm existence of the product / feature but exact throughput and pricing are not in this pass. Verify specific capability claim against the underlying vendor docs and the actual license / pricing matrix before stating it on the record; do not paraphrase per-platform pricing or license terms into specific dollar figures or commercial-use clauses.
- MLX docs explicitly position the framework for Apple Silicon with unified memory. Verify specific capability claim against the underlying vendor docs and the actual license / pricing matrix before stating it on the record; do not paraphrase per-platform pricing or license terms into specific dollar figures or commercial-use clauses.
- Docs confirm framework support but the specific model coverage is not in this pass. Verify specific capability claim against the underlying vendor docs and the actual license / pricing matrix before stating it on the record; do not paraphrase per-platform pricing or license terms into specific dollar figures or commercial-use clauses.
演示思路
- 六个 vendor 同 token 流延迟测量并排 —— 同一模型、同一 prompt,画 time-to-first-token 和 token 间延迟
- 决策树:「哪个推理 vendor 配哪个用例」(低延迟流式 → Groq LPU,单芯片大模型 → Cerebras WSE,设备端 / 隐私 → Apple MLX,AWS 原生 → Inferentia3,Google Cloud → TPU v6,裸 GPU → Blackwell)
- 成本计算器:chat 工作负载下六个 vendor per-1K-token 成本对比