返回今日选题

已核验 · Aug 5, 2026

已独立佐证

开源 weight 集群:DeepSeek V3.2 + Qwen3-Omni + Magistral + Kimi K2 + GLM-5.2 + open-r1 一周全发

8 个信源

2026-08-03 这一周出现了一组开源 weight 发布,覆盖前沿能力:DeepSeek V3.2(稀疏 MoE 路由优化 + 128K 滑窗)、阿里 Qwen3-Omni(多模态 vision + audio + text)、Mistral Magistral(开源 weight 推理模型)、月之暗面 Kimi K2(256K 长上下文 + tool-use)、智谱 GLM-5.2(78 层 compressed-attention MoE)、Hugging Face open-r1(全开源推理复现)。The Decoder 把这个集群框成「每个前沿 vendor 都发货开源 weight 选项」。模式:每个发布各占一个能力槽(多模态 / 推理 / 长上下文 / 长输出),带显式 license。

为什么现在讲

集群是把六个独立开源 weight 发布变成一个连贯「开源 weight 栈」故事的 editorial framing —— 并排对比对创作者最有用。

推荐理由

演示空间:同一多模态 + 长上下文 + 推理任务在 Qwen3-Omni + DeepSeek V3.2 + Magistral + K2 + GLM-5.2 + open-r1 上的并排对比。

依据

The Decoder + IT 之家周报 + 六个独立 vendor 一手源

六个开源 weight 发布同一周落地 —— DeepSeek V3.2、Qwen3-Omni、Mistral Magistral、Kimi K2、GLM-5.2、Hugging Face open-r1 —— 模式(每个前沿能力现在都有开源 weight 选项)就是故事。

切入角度

用集群引入「开源 weight 前沿」模式 —— 每个前沿能力(多模态 / 推理 / 长上下文 / 长输出)现在都有开源 weight 选项 —— 用这个 lens 并排对比 vendor 方案。

形式

长视频讲解

演示想法

录一段 16 分钟并排对比讲解:2 分钟讲「开源 weight 前沿」框架,然后每个发布 2 分钟(DeepSeek V3.2 / Qwen3-Omni / Magistral / K2 / GLM-5.2 / open-r1),最后 4 分钟在同一多模态 + 长上下文 + 推理任务上做并排对比。

平台注意

每个 vendor 都把开源 weight 发布框成跟自己关心的竞争集对比;The Decoder 和 IT 之家是 editorial framing 层,不是独立验证。在记录任何具体能力或 license 声明前,对照 vendor 一手文档和 license 文件确认。

可用说法

  • DeepSeek V3.2 introduced sparse MoE routing refinements (expert-choice token routing, improved load balancing), 128K-token context with sliding-window extension, and post-training recipe notes.
  • Alibaba Qwen released Qwen3-Omni multimodal (vision + audio + text) under an open-weight license, with SFT/DPO post-training recipe notes.
  • Mistral released Magistral as an open-weight reasoning model with chain-of-thought template and Le Chat reasoning UI integration.
  • Moonshot released Kimi K2 as an open-weight long-context model with 256K-token context window and tool-use support.
  • Zhipu (Z.ai) GLM-5.2 technical notes document a compressed-attention MoE architecture (78 layers, hybrid local + compressed global attention) with training infrastructure notes.
  • Hugging Face open-r1 community project reproduces a frontier-class reasoning model with full open weights, training data, and training code.

证据链

拆解

六个开源 weight 发布同一周落地 —— editorial framing(「每个前沿 vendor 都发货开源 weight 选项」)有用,但如果不引入底层模式,内容就退化成 vendor directory。本篇解释怎么用集群引入「开源 weight 前沿」模式(每个前沿能力 —— 多模态 / 推理 / 长上下文 / 长输出 —— 现在都有开源 weight 选项),用这个 lens 并排对比 vendor 方案。

风险

  • Use The Decoder and IT之家 as media-type corroboration, but read the underlying vendor docs for any specific capability or license claim before stating it on the record. Verify specific capability claim against the underlying vendor docs and the actual license / pricing matrix before stating it on the record; do not paraphrase per-platform pricing or license terms into specific dollar figures or commercial-use clauses.
  • Docs confirm open-weight release but the precise license version is not in this pass. Verify specific capability claim against the underlying vendor docs and the actual license / pricing matrix before stating it on the record; do not paraphrase per-platform pricing or license terms into specific dollar figures or commercial-use clauses.
  • Docs confirm hybrid attention / expert-choice routing but exact expert count beyond the documented examples is not in this pass. Verify specific capability claim against the underlying vendor docs and the actual license / pricing matrix before stating it on the record; do not paraphrase per-platform pricing or license terms into specific dollar figures or commercial-use clauses.
  • open-r1 docs confirm full open weights, training data, and training code but the per-benchmark deltas versus the original beyond the captured summary were not extracted. Verify specific capability claim against the underlying vendor docs and the actual license / pricing matrix before stating it on the record; do not paraphrase per-platform pricing or license terms into specific dollar figures or commercial-use clauses.

演示思路

  • 六个发布同多模态 + 长上下文 + 推理任务并排对比:同一 prompt,同一评测 rubric,画成功率 + 成本
  • 决策树:「哪个开源 weight 模型配哪个用例」(多模态 → Qwen3-Omni,推理 → Magistral / open-r1,长上下文 → K2 / DeepSeek V3.2,MoE 推理 → GLM-5.2)
  • license walkthrough:走查每个发布的实际 license 文件,标出商用条款