返回今日选题

已核验 · Aug 5, 2026

已独立佐证

Qwen3-Omni + DeepSeek V3.2:多模态 + 长上下文开源 weight pair

3 个信源

阿里 Qwen3-Omni(开源 weight 多模态:vision + audio + text 带 SFT/DPO post-training)+ DeepSeek V3.2(稀疏 MoE 路由优化 + 128K 滑窗上下文扩展),创作者拿到一个开源 weight 多模态 + 长上下文 pair。两个发布各占不同能力轴 —— Qwen3-Omni 占多模态,DeepSeek V3.2 占长上下文 —— 合起来让创作者发一个能处理 100K+ token 输入的多模态 agent,不用付前沿 vendor 价。

为什么现在讲

两个发布同一周落地 —— 开源 weight 多模态 + 长上下文 pair 现在两件可买。

推荐理由

演示空间:live 多模态 + 长上下文 agent,摄取 90K token 转录文本并回答关于它的、附视觉 grounding 的问题。

依据

Qwen readthedocs + DeepSeek API 文档 + The Decoder 周报

Qwen3-Omni 和 DeepSeek V3.2 本周一起发 —— 两者合起来让你发一个能处理 100K+ token 输入的多模态 agent,不用付前沿 vendor 价。

切入角度

用 Qwen3-Omni + DeepSeek V3.2 pair 引入「开源 weight 多模态 + 长上下文」模式 —— 两个开源 weight 发布各占互补能力轴。

形式

长视频讲解

演示想法

录一段 10 分钟讲解:2 分钟讲「开源 weight 多模态 + 长上下文」pair 框架,3 分钟讲 Qwen3-Omni(多模态),3 分钟讲 DeepSeek V3.2(长上下文),2 分钟做 live 多模态 + 长上下文 agent demo。

平台注意

两者精确开源 weight license 条款本次未抽取,商用前对照实际 license 文件确认。MoE 架构细节(expert 数量、路由优化)只文档化到高层。

可用说法

  • Alibaba Qwen released Qwen3-Omni multimodal (vision + audio + text) under an open-weight license, with SFT/DPO post-training recipe notes.
  • DeepSeek V3.2 introduced sparse MoE routing refinements (expert-choice token routing, improved load balancing), 128K-token context with sliding-window extension, and post-training recipe notes.

证据链

拆解

Qwen3-Omni 和 DeepSeek V3.2 各占不同能力轴 —— Qwen3-Omni 占多模态,DeepSeek V3.2 占长上下文。本篇把两者框成「pair」一起覆盖多模态 + 长上下文,同时保持逐发布差异诚实(Qwen3-Omni 是多模态 vision + audio + text;DeepSeek V3.2 是稀疏 MoE + 128K 滑窗)。

风险

  • Docs confirm open-weight release but the precise license version is not in this pass. Verify specific capability claim against the underlying vendor docs and the actual license / pricing matrix before stating it on the record; do not paraphrase per-platform pricing or license terms into specific dollar figures or commercial-use clauses.
  • Use The Decoder and IT之家 as media-type corroboration, but read the underlying vendor docs for any specific capability or license claim before stating it on the record. Verify specific capability claim against the underlying vendor docs and the actual license / pricing matrix before stating it on the record; do not paraphrase per-platform pricing or license terms into specific dollar figures or commercial-use clauses.
  • Docs confirm hybrid attention / expert-choice routing but exact expert count beyond the documented examples is not in this pass. Verify specific capability claim against the underlying vendor docs and the actual license / pricing matrix before stating it on the record; do not paraphrase per-platform pricing or license terms into specific dollar figures or commercial-use clauses.

演示思路

  • live 多模态 + 长上下文 agent:摄取 90K token 视频转录文本,用 Qwen3-Omni 的视觉 grounding 回答问题
  • 成本对比:开源 weight 多模态 + 长上下文 pair vs 闭源前沿(Claude / GPT)—— 画 per-1K-token 成本
  • 路由 demo:多模态查询路由到 Qwen3-Omni,长上下文查询路由到 DeepSeek V3.2 —— 量延迟和成本