已核验 · Aug 5, 2026
已独立佐证Blackwell + Neuron SDK v3 + TPU v6:Blackwell-era 云推理栈 —— 三个 vendor,三个对同一 GPU 代的答案
4 个信源NVIDIA Blackwell(定义这一时期的 GPU 代)、AWS Inferentia3 + Neuron SDK v3(AWS 对 Blackwell-era 推理的答案)、Google TPU v6 Trillium(Google 的答案)合起来定义 Blackwell-era 云推理栈。每个 vendor 在 cost / latency / model coverage 上定位不同 —— NVIDIA 在裸 GPU 吞吐 + 生态锁定,AWS 在单实例定价 + Neuron-native 模型覆盖,Google 在 TPU pod 拓扑 + Vertex AI 集成。三者并不等价:部署模型不同,SDK 不同,逐模型性能特征不同。
为什么现在讲
三个同一窗口内 GA —— Blackwell-era 云推理现在三种形状可买。
推荐理由
演示空间:同一生产推理工作负载部署在 Blackwell、Inferentia3、TPU v6 上,量延迟 / 吞吐 / 成本。
依据
三个独立 vendor 一手源 + The Decoder 周报
“Blackwell、Inferentia3、TPU v6 同一窗口内全发 —— 三个 Blackwell-era 云推理答案、三种不同的 cost、latency、model coverage 配方 —— 选择对创作者工作负载很重要。”
切入角度
用三个 Blackwell-era 答案引入「云推理栈」框架 —— 三个 vendor、三种不同的 cost / latency / model coverage / 生态 配方 —— 用这个 lens 讨论哪个配哪个创作者工作负载。
形式
长视频讲解
演示想法
录一段 12 分钟讲解:3 分钟讲「Blackwell-era 云推理栈」框架,3 分钟讲每个 vendor(Blackwell / Inferentia3 / TPU v6),3 分钟做 live 并排 demo,同一生产推理工作负载在三个上跑。
平台注意
逐模型推理吞吐数字和单实例定价矩阵本次未抽取,不要给出具体 tokens-per-second 或 dollar-per-hour 数字。Neuron SDK 支持 PyTorch / JAX / TensorFlow,但具体前沿模型覆盖(超出本次捕获范围)未抽取。
可用说法
- NVIDIA Blackwell generation (B200 datasheet) documents FP4/FP8 inference throughput, NVLink Switch fabric scaling, and DGX SuperPOD reference architecture for production inference.
- AWS Inferentia3 GA + Neuron SDK v3 with PyTorch / JAX / TensorFlow support, distributed inference, and the Neuron Compiler for custom model compilation.
- Google Cloud TPU v6 (Trillium) GA on Cloud with Vertex AI managed inference integration and per-pod scaling topology.
证据链
来自新闻
拆解
Blackwell、Inferentia3、TPU v6 是三个不同的 Blackwell-era 云推理答案,每个 cost / latency / model coverage / 生态 配方不同。本篇让三个 vendor 差异保持诚实(部署模型、SDK、逐模型性能特征),而不是压缩成「Blackwell = 快」的单线叙事。
信源
风险
- Vendor docs confirm existence of the product / feature but exact throughput and pricing are not in this pass. Verify specific capability claim against the underlying vendor docs and the actual license / pricing matrix before stating it on the record; do not paraphrase per-platform pricing or license terms into specific dollar figures or commercial-use clauses.
- Docs confirm framework support but the specific model coverage is not in this pass. Verify specific capability claim against the underlying vendor docs and the actual license / pricing matrix before stating it on the record; do not paraphrase per-platform pricing or license terms into specific dollar figures or commercial-use clauses.
- Use The Decoder and IT之家 as media-type corroboration, but read the underlying vendor docs for any specific throughput or pricing claim before stating it on the record. Verify specific capability claim against the underlying vendor docs and the actual license / pricing matrix before stating it on the record; do not paraphrase per-platform pricing or license terms into specific dollar figures or commercial-use clauses.
演示思路
- live 生产推理工作负载:同一 chat 工作负载在 Blackwell / Inferentia3 / TPU v6 上,量延迟 / 吞吐 / 成本
- vendor 定位矩阵:把每个 vendor 画在 cost vs latency vs model coverage 上
- 迁移故事:「把工作负载从 Blackwell 迁到 Inferentia3」,量 porting 工作量和成本 delta