Aug 5, 2026
官方
Anthropic: Constitutional AI v3 specification
Anthropic 发布 Constitutional AI v3 规范
Anthropic Constitutional AI v3 文档化多阶段训练(显式原则、无害 reward model + 有用 reward model)、公开的批评-修订循环。
今日资讯
面向创作者调研的 AI 资讯流。资讯在完成核验和选题包装前,会和今日选题保持分离。
Aug 5, 2026
官方
Anthropic: Constitutional AI v3 specification
Anthropic Constitutional AI v3 文档化多阶段训练(显式原则、无害 reward model + 有用 reward model)、公开的批评-修订循环。
Aug 5, 2026
官方
OpenAI: Preparedness Framework v2 update
OpenAI Preparedness Framework v2 文档化风险类别、分数阈值、缓解义务、跨职能审查委员会。
Aug 5, 2026
官方
Google DeepMind: Frontier Safety Framework v3
Google DeepMind Frontier Safety Framework v3 文档化能力评估、早期预警指标、缓解部署义务。
Aug 5, 2026
官方
Meta: Llama Guard 4 open-weight safety classifier
Llama Guard 4 model card 文档化开源 weight 安全分类器、不安全内容的多类 taxonomy、与 Llama Stack 1.5 reference server 集成。
Aug 5, 2026
官方
Microsoft: AI Red Team + AI Risk Shadow Model updates
Microsoft AI Red Team 文档化 AI Red Team 方法论更新、AI Risk Shadow Model 框架(用于内部 red-team 跟踪)。
Aug 5, 2026
官方
C2PA: Content Credentials 2.0 specification
C2PA Content Credentials 2.0 文档化媒体资产的密码学 provenance、模型 + 工具 attestations、抗篡改 manifest。
Aug 5, 2026
媒体
The Decoder: tech-press coverage of the 'safety frameworks' cluster for the week of 2026-08-05
The Decoder 把 2026-08-05 安全与治理发布集群(Constitutional AI v3、Preparedness v2、Frontier Safety v3、Llama Guard 4、AI Red Team、C2PA 2.0)框成「每个前沿 lab 都发货安全框架更新」。
Aug 5, 2026
媒体
IT之家: 中文科技媒体覆盖 8/5 安全框架集群
IT 之家首页快照覆盖 2026-08-05:Anthropic Constitutional AI 中文对齐、OpenAI Preparedness 中文社区、Google DeepMind Frontier Safety 中文研究、Meta Llama Guard 4 中文开源、Microsoft AI Red Team 中文企业、C2PA Content Credentials 中文内容真实性。
Aug 5, 2026
官方
C2PA: Content Credentials 2.0 specification
把 Llama Guard 4(开源 weight 安全分类器,做输入/输出过滤)+ C2PA Content Credentials 2.0(媒体资产密码学 provenance)组合,创作者拿到一个能实际部署在自己平台的栈。
Aug 5, 2026
官方
OpenAI: Preparedness Framework v2 update
OpenAI Preparedness v2(风险类别、分数阈值、缓解义务)+ Google DeepMind Frontier Safety v3(能力评估、早期预警、缓解部署义务)定义前沿 lab 公布的双评估框架。
Aug 5, 2026
官方
NVIDIA: Blackwell B200 GPU datasheet and inference benchmarks
NVIDIA 数据中心 GPU 产品页文档化 Blackwell 代 B200:FP4/FP8 推理吞吐、NVLink Switch 互联扩展、DGX SuperPOD 参考架构。
Aug 5, 2026
官方
Groq: LPU inference engine v3 — deterministic token-stream latency
Groq 首页文档化 LPU 推理引擎 v3:确定性 token 流延迟(可预测的 per-token time-to-next-token,与 batch size 无关)、LPU 机架级部署。
Aug 5, 2026
官方
Cerebras: WSE wafer-scale inference + cloud SDK
Cerebras 首页文档化 WSE 晶圆级推理 + cloud SDK —— 单芯片推理,避免跨 GPU 模型并行。
Aug 5, 2026
官方
Apple: MLX 1.0 framework + MLX-LM for on-device inference
Apple MLX 文档化 MLX 1.0 framework GA、MLX-LM Python 包(Apple Silicon 设备端 LLM 推理)、统一内存模型。
Aug 5, 2026
官方
AWS: Inferentia3 + Neuron SDK v3 GA for production inference
AWS Neuron SDK 文档化 Inferentia3 GA、Neuron SDK v3 支持 PyTorch / JAX / TensorFlow、分布式推理、Neuron Compiler 自定义模型编译。
Aug 5, 2026
官方
Google Cloud: TPU v6 (Trillium) GA on Cloud + Vertex AI integration
Google Cloud TPU 文档化 TPU v6 (Trillium) Cloud GA + Vertex AI 托管推理集成 + per-pod 扩展拓扑。
Aug 5, 2026
媒体
The Decoder: tech-press coverage of the 'inference hardware' cluster for the week of 2026-08-04
The Decoder 把 2026-08-04 推理硬件 / 框架发布集群框成「每个推理 vendor 都发货一个 Blackwell-era 答案」。
Aug 5, 2026
媒体
IT之家: 中文科技媒体覆盖 8/4 推理硬件集群
IT 之家首页快照覆盖 2026-08-04:NVIDIA Blackwell 中文数据中心、Groq LPU 中文开发者、Cerebras WSE 中文落地、Apple MLX 1.0 苹果生态、AWS Inferentia3 中文 Neuron SDK、Google TPU v6 中文 Vertex AI。
Aug 5, 2026
官方
Apple: MLX 1.0 framework + MLX-LM for on-device inference
把 Apple MLX 1.0(framework GA)+ MLX-LM(设备端 LLM 推理 Python 包)组合,创作者拿到一个 Apple Silicon 设备端推理栈 —— 适合隐私敏感 + 离线可用场景。
Aug 5, 2026
官方
Google Cloud: TPU v6 (Trillium) GA on Cloud + Vertex AI integration
NVIDIA Blackwell(GPU 代)、AWS Inferentia3 + Neuron SDK v3(AWS 答案)、Google TPU v6(Google 答案)定义 Blackwell-era 云推理栈。每个 vendor 的答案在 cost / latency / model coverage 上定位不同。
Aug 5, 2026
官方
DeepSeek: V3.2 technical report — sparse MoE routing refinements
DeepSeek API 文档新闻条目文档化 V3.2 稀疏 MoE 路由优化(expert-choice token routing、改进 load balancing)、128K token 滑窗 attention 扩展、post-training 配方说明。
Aug 5, 2026
官方
Alibaba Qwen: Qwen3-Omni open-weight multimodal release on Hugging Face
Qwen readthedocs 文档化 Qwen3-Omni 多模态(vision + audio + text)采用开源 weight 协议,带 SFT/DPO post-training 配方说明。
Aug 5, 2026
官方
Mistral: Magistral reasoning model open-weight release
Mistral reasoning 文档文档化 Magistral 作为开源 weight 推理模型,带 chain-of-thought 模板和 Le Chat reasoning UI 集成。
Aug 5, 2026
官方
Moonshot Kimi: K2 long-context open-weight release
Moonshot Kimi 平台文档文档化 K2 开源 weight 长上下文发布,256K token 上下文窗口,tool-use 支持。
Aug 5, 2026
官方
Zhipu (Z.ai) GLM: GLM-5.2 compressed-attention MoE technical notes
Z.ai GLM 模型文档文档化 GLM-5.2 compressed-attention MoE 架构(78 层,hybrid local + compressed global attention)带训练基础设施说明。
Aug 5, 2026
官方
Hugging Face: open-r1 fully open reasoning model reproduction
Hugging Face open-r1 社区项目复现前沿级推理模型,全开源 weight + 训练数据 + 训练代码。
Aug 5, 2026
媒体
The Decoder: tech-press coverage of the 'open-weight release' cluster for the week of 2026-08-03
The Decoder 把 2026-08-03 开源 weight 集群(DeepSeek V3.2、Qwen3-Omni、Magistral、K2、GLM-5.2、open-r1)框成「每个前沿 vendor 都发货开源 weight 选项」。
Aug 5, 2026
媒体
IT之家: 中文科技媒体覆盖 8/3 开源 weight 集群
IT 之家首页快照覆盖 2026-08-03:DeepSeek V3.2 国内适配、Qwen3-Omni 多模态开源、Mistral Magistral 中文文档、Kimi K2 长上下文中文工具链、GLM-5.2 国内生态、HF open-r1 中文社区复现。
Aug 5, 2026
官方
Mistral: Magistral reasoning model open-weight release
把 Mistral Magistral(开源 weight 推理模型带 chain-of-thought 模板)+ Hugging Face open-r1(全开源复现前沿级推理模型)组合,创作者就拿到一个开源 weight 推理栈。
Aug 5, 2026
官方
Alibaba Qwen: Qwen3-Omni open-weight multimodal release on Hugging Face
Qwen3-Omni(多模态 vision + audio + text)+ DeepSeek V3.2(长上下文 128K 带滑窗扩展),创作者拿到一个开源 weight 多模态 + 长上下文 pair。
Aug 5, 2026
媒体
The Decoder: weekly roundup of model and tooling releases for the week of 2026-07-27 — 2026-08-02
The Decoder 周报首页快照覆盖 2026-07-27 — 2026-08-02 周,聚合十余条 vendor 发布(OpenAI、Anthropic、Google、Microsoft、Meta、NVIDIA),归入「tools consolidation」框架。
Aug 5, 2026
媒体
IT之家: 2026-08-01 weekly AI vendor news roundup
IT 之家首页快照覆盖 2026-08-01 周:Anthropic Claude 4.5 Sonnet 中文适配、OpenAI Realtime 中文开发者覆盖、Microsoft Copilot Studio 企业落地、Meta Llama 4 国内开源生态、NVIDIA NIM 中文文档。
Aug 5, 2026
官方
OpenAI: ChatGPT Scheduled Tasks GA + Tasks API
OpenAI 平台文档化 Scheduled Tasks 正式 GA(创建 recurring 或 one-shot task,对存好的 prompt 触发模型调用),并提供 Tasks API 用于编程式管理。
Aug 5, 2026
官方
Anthropic: Claude Skills GA — procedural skill bundles for repeated workflows
Anthropic Skills 文档化 Claude Skills 正式 GA:文件夹打包的程序化知识(SKILL.md + 脚本 + references),通过 tool_search 按需加载,模型判断工作流需要时自动调用。
Aug 5, 2026
官方
Vercel: AI Gateway GA — unified LLM routing, fallbacks, cost analytics
Vercel AI Gateway 文档化 GA:统一路由前置多模型 provider,带 fallback、retry、spend limits、按请求成本归因;gateway 处理 auth token 轮换和跨 provider streaming。
Aug 5, 2026
官方
Notion: Notion AI Agents GA — autonomous workflow agents in the workspace
Notion 帮助文档化 Notion AI Agents GA:工作区常驻 agent,可以读、写并触发 Notion 工作流;agent 限定到特定 page 和 database,通过 trigger(手动、定时、事件)运行。
Aug 5, 2026
官方
Linear: Linear AI for issue triage — automated labeling and routing
Linear AI 文档化 AI 驱动的 issue triage —— 自动化 project / label / assignee 推荐、重复检测、AI 生成 summary;定位为「首轮 triage,浮现而不替代人审」。
Aug 5, 2026
官方
Figma: Figma Make GA — prompt-to-design prototype generator
Figma 帮助文档化 Figma Make GA:prompt-to-design 原型生成器,从自然语言描述产出可编辑 Figma primitives(frames / components / auto-layout),不是栅格图。
Aug 5, 2026
媒体
The Decoder: tech-press coverage of the 'productivity agents' GA cluster for the week of 2026-08-02
The Decoder 把 2026-08-02 GA 集群(OpenAI Scheduled Tasks、Anthropic Skills、Vercel AI Gateway、Notion AI Agents、Linear AI triage、Figma Make)框成「每个 productivity tool 都发货 agent 面」。
Aug 5, 2026
媒体
IT之家: 中文科技媒体覆盖 8/2 productivity agents GA 集群
IT 之家首页快照覆盖 2026-08-02:ChatGPT Scheduled Tasks 中文上手、Anthropic Skills 中文文档、Vercel AI Gateway 中文落地、Notion AI Agents 中文模板、Linear AI triage 中文项目管理、Figma Make 国内设计社区试用。
Aug 5, 2026
官方
Vercel: AI Gateway GA — unified LLM routing, fallbacks, cost analytics
把 OpenAI Scheduled Tasks(recurring 模型调用)、Anthropic Claude Skills(按需程序化工作流)、Vercel AI Gateway(多 provider 路由)组合,创作者就能拿到一个 always-on agent runtime,不用自建基础设施。
Aug 5, 2026
官方
Notion: Notion AI Agents GA — autonomous workflow agents in the workspace
Figma Make(设计 primitives)、Notion AI Agents(工作区工作流)、Linear AI(issue triage)共享一个模式:productivity tool 发货一个 agent,住在 tool 现有数据模型里,scope 到 tool 的 primitives(frames / pages / issues)。
Aug 5, 2026
官方
OpenAI: GPT Realtime GA + Realtime pricing tier for voice agents
OpenAI 的 GPT Realtime 模型页文档化了生产 GA 档位,带服务端 VAD、session 中途 system instruction 更新、live session 中的 function calling;另有 Realtime-mini 档位定位 always-on 语音 agent。
Aug 5, 2026
官方
Anthropic: Claude 4.5 Sonnet tools release notes (programmatic tool calling + web fetch GA)
Anthropic Claude 4.5 Sonnet release notes 文档化了三项 GA:程序化 tool calling(在 code execution 沙盒里定义并调用 tool)、web_fetch server tool(带 caching 和 provenance metadata)、1M token 上下文 GA。
Aug 5, 2026
官方
Google Cloud Vertex AI: Gemini 2.5 Pro deep-research mode and grounding updates
Google Cloud Vertex AI Model Garden 文档化 Gemini 2.5 Pro deep-research 模式(长程 agentic research,带引用)、Grounding with Google Search 默认启用、JSON schema 结构化输出、File Search 检索工具。
Aug 5, 2026
官方
Meta AI: Llama 4 multimodal checkpoints and Llama Stack 1.5 reference server
Meta Llama 4 model card 文档化 vision+audio 多模态 checkpoint 变体、开源 Llama Stack 1.5 参考服务(推理 + 内置 safety guardrails)、SFT/DPO post-training 配方说明。
Aug 5, 2026
官方
NVIDIA NIM: catalog updates for open-weight LLMs and agent runtimes
NVIDIA NIM for LLMs 文档化近期开源 weight checkpoint 上架、catalog 全面支持 tool/function calling、接入 NeMo Agent toolkit runtime orchestration。
Aug 5, 2026
官方
Microsoft: Copilot Studio GA for autonomous agent workflows and Teams distribution
Microsoft Learn Copilot Studio 文档化 multi-step autonomous agent workflow GA、原生 Teams 分发、Power Automate action 集成、actions-pack 计费档位。
Aug 5, 2026
官方
Hugging Face TRL: PPO/GRPO/DPO trainer updates and on-policy distillation recipe
TRL 库文档化统一 GRPOConfig(跨 PPO/GRPO/DPO trainer)、对齐 SEED/LongStraw 论文方向的 on-policy distillation 配方、PPO trainer 加强 LoRA 支持。
Aug 5, 2026
官方
GitHub: GitHub Models catalog and Copilot Chat multi-model switcher
GitHub Marketplace models catalog 文档化扩展的前沿 + 开源 weight 模型 catalog、Copilot Chat 多模型切换器在 IDE 上线、free-tier playground 按请求 token 配额。
Aug 5, 2026
论文
arXiv 2607.14777: SEED — Self-Evolving On-Policy Distillation for Agentic Reinforcement Learning
Hugging Face 2026-07-17 索引的 arXiv 2607.14777 提出 SEED —— 一个用于长周期 agentic RL 的自演化 on-policy distillation 框架,把已完成轨迹转成 hindsight skills 再蒸馏回策略,消除稀疏奖励的 credit assignment 问题;两阶段:(1) Hindsight Skill SFT 用外部 analyzer(GLM-5.2)对 1,440 条轨迹(180 任务 × K0=8 rollouts)做注解,抽取可复用的自然语言 skills;(2) Self-Evolving OPD 当前 policy snapshot 同时担任 rollout actor 与 trajectory analyzer,每轮刷新让 actor 与 analyzer 共同演化;同一模型 actor + analyzer 共享参数;confidence-gated token-level distillation;sigmoid(β_opd × Δlog-prob);联合损失 L_SEED = L_RL(GRPO + KL 正则)+ λ_opd · L_OPD;梯度只流过 ordinary student 分支;推理时策略只从 ordinary history 行动。自报基准(ALFWorld / Search-QA / WebShop score / WebShop succ):Qwen2.5-3B-Instruct GRPO 75.0/36.4/79.8/63.3 vs Seed 91.8/45.7/88.5/78.9;Qwen2.5-7B-Instruct GRPO 81.2/42.0/80.9/72.6 vs Seed 96.1/48.6/89.7/78.1;Qwen3-1.7B-Instruct GRPO 46.1/40.8/67.3/38.3 vs Seed 92.0/42.2/87.1/77.3;sample efficiency:Seed 用 60% 训练数据(ALFWorld 80.7)就超过全量 GRPO(75.0);cross-domain:ALFWorld Unseen split +15.3 分(86.2 vs 70.9);multimodal:Qwen2.5-VL-3B Sokoban 82.0% / EZPoints 100.0%,平均 91.0% vs GRPO 77.0%。
Aug 5, 2026
论文
arXiv 2607.14952: LongStraw — Long-Context RL Beyond 2M Tokens under a Fixed GPU Budget
Hugging Face 2026-07-17 索引(207 赞,#1 Daily Paper)的 arXiv 2607.14952(Mind Lab)提出 LongStraw —— 一个固定 GPU 预算下做百万 token RL post-training 的架构感知 execution stack,以 GRPO 实例化;方法:只用一次无 autograd 评估 shared prompt,只保留后续 token 所需的 model-specific state,短 response 分支在 autograd 下逐个 replay,把 live training graph 从「整段 prompt + response」压到「单条 response 分支」,用 replay 时间换 GPU 内存;模型实现:Qwen3.6-27B(hybrid recurrent + full-attention)与 GLM-5.2(compressed-attention MoE,78 层)。自报数字:8x H20 GPU 上 group=2 与 group=8 完成 grouped Qwen scoring + response backward 在 2.1M positions;group size 增长每档仅 +0.21 GB peak allocated memory;压力测试达到 4.46M positions;32x H20 GPU 上对 2.1M-token prompt 跑完整 78 层 GLM-5.2 做端到端验证。关键洞察:瞄准「inference(~1M token 上下文)与 RL post-training(经常 ≤256K token)」之间的差距,对带累积轨迹的 AI agent 尤其相关;作者说明这些实验只建立 execution capacity、不构成完整训练正确性。
Aug 5, 2026
论文
arXiv 2607.14935: VideoChat3 — Fully Open Video MLLM for Efficient and Generalist Video Understanding
Hugging Face 2026-07-17 索引的 arXiv 2607.14935(MCG-Nanjing University)提出 VideoChat3 —— 4B 参数全栈开放视频 MLLM;架构:(1) Inflated 3D Vision Transformer (I3D-ViT)把 2D 空间 self-attention 扩成 3D 时空 self-attention,16x 时空压缩比;(2) Adaptive Frame Resolution for Streaming Video Perception 用 </Silence> / </Standby> / </Response> 三个状态 token 控制下一窗口像素预算(Silence/Response 224²,Standby 448²);训练数据 3M 样本:VideoChat3-Academic2M(2.27M)、VideoChat3-LV116K(116.2K)、VideoChat3-OL617K(617K),四阶段训练(tokenizer pre-training → video-language alignment → video instruction tuning → long & streaming instruction tuning);全栈开放 —— 权重 + 训练代码 + 训练策略 + 完整训练数据。自报基准(VideoChat3-4B vs open-weight Qwen3-VL-4B):MotionBench 61.7 vs 58.6、TempCompass 75.6 vs 70.8、Video-MME 70.1 vs 69.3、LVBench 56.7 vs 56.2、MMVU 56.4 vs 50.5、Charades TL mIoU 56.1 vs 46.4、VUE-TR V1 47.9 vs 32.9、VUE-TR V2 40.2 vs 19.6、MomentSeeker 25.9 vs 13.8;流式 ODVBench 72.3 vs StreamForest 59.9(+12.4);NVIDIA H200 上 2048 帧:总延迟 20.412s vs Qwen3-VL 44.449s、视觉 token 100,352 vs 200,704(一半)、GPU 内存省 26.14 GB;文章称在三个 TimeLens split 上超过 GPT-5 与 Gemini 2.5 Flash。
Aug 5, 2026
社区
IBM Research blog: It's time for cryptography to get its own abstraction layer
IBM Research 2026-07-17 博客《It's time for cryptography to get its own abstraction layer》主张加密学需要一层把高级意图与底层实现分开的标准层(参考当年文件系统与 socket 对存储与网络做的事);IBM 提议一个按 scopes 组织的 intent-based API,每个 scope 代表一类加密意图(标准数字签名 / 认证加密等),应用只表达「需要什么」、算法/参数/实现下沉到下层集中管理;关键特性:把策略(控制平面)与执行(数据平面)分开,参考 SDN 对网络的建模;把加密后端当作「单一接口后可互换的 provider」覆盖软件/硬件/云/TEE;允许算法升级不动应用代码;支持 PQC 迁移不必「替换算法本身」;IBM 同步放出 API spec、参考 standalone server、Go client SDK,邀请社区「explore the work, challenge the assumptions, help shape its evolution」。
Aug 5, 2026
官方
Cursor changelog: Improvements to Cursor in Slack
Cursor 2026-07-17 changelog 记录三项改进:(1) 交互改进 —— Cursor 在开始前先回一个 plan,用户可早改向;执行中更新状态;in-message buttons 替换为「compact footer links」;表格 / PRs / 制品渲染更干净;(2) Multi-repo 环境支持 —— Cursor 可以在命名的多 repo 环境中启动(而非单一默认仓库),目标环境能访问所有相关 repo;运行中可用「Switch repository button」加进当前环境之外的 repo,Cursor 从中断处继续;(3) Cross-channel workflows —— Cursor 可在多个 Slack 频道/thread 之间读写,把上下文从工作区其他地方拉进来,把更新发回原 thread 或相关频道。
Aug 5, 2026
论文
arXiv 2607.14749: WanSong v1.0 Technical Report
Hugging Face 2026-07-17 索引(arXiv 提交 7/16)的 arXiv 2607.14749(Wan-AI)给 WanSong v1.0 Technical Report —— 商用级 text-to-music 基础模型,单次端到端生成最长约 5 分钟歌曲,输出可分离的 vocal 与 BGM 双 stem;架构:连续 1-D VAE(44.1 kHz → 64 通道潜在,43.1 Hz 下采样 1024 倍)对抗式训练 + hybrid-MMDiT transformer(~25B 参数)处理 LLM 文本 token 与 dual-stem VAE 音频 token;能力:多语言(中/英/日/韩)最长 5 分钟歌曲、3 阶段预训练(90s → 300s → SFT)、为推理加速的 step-distillation、用 DPO 再 ReFL 的 RLHF 对齐;自报基准(WanSong Bench,~200 条 4 分钟 clips):PER 7.43% vs Suno V5 22.80% / Mureka V7.6 12.7% / LeVo 27.11%;SongEval 4.47/4.55/4.57/4.46/4.4;Muq 文本对齐 0.44;内部 musicality 5.49 vs Suno V5 4.18 / Mureka V7.6 3.83 / LeVo 1.69。
Aug 5, 2026
论文
Hugging Face: SEED ablation thread
SEED 7/17 论文 ablation(ALFWorld avg):去掉 Hindsight-Skill SFT → 86.0(-5.8);去掉 Self-Evolving OPD → 87.0(-4.8);用静态 offline skills 替代 on-policy skills → 84.4(-7.4);三段 ablation 加起来约 -18 分,意味着 on-policy 自演化与 hindsight-skill SFT 是两个独立的贡献项,各贡献一半左右;用 60% 训练数据(ALFWorld 80.7)就超过全量 GRPO(75.0)展示 SEED sample efficiency 显著高于 baseline。
Aug 5, 2026
论文
Hugging Face: LongStraw cluster thread
LongStraw 7/17 实现覆盖三个模型:Qwen3.6-27B(hybrid recurrent + full-attention)、GLM-5.2(compressed-attention MoE,78 层);HF daily papers 把 LongStraw 列为 7/17 #1 Daily Paper,获 207 赞;GitHub MindLab-Research/longstraw 41 stars;社区评论 O96a 质疑:对真实 agent loop(上下文在轨迹中间不可预测地增长)而非干净切成 prompt + generation 的场景,适用性如何?
Aug 5, 2026
媒体
The Decoder: IBM Research cryptography abstraction layer
The Decoder 2026-07-17 tech-press 报道 IBM Research 加密学抽象层提议(intent-based API、scopes 代表加密意图、策略/执行分离借鉴 SDN、provider 可互换、PQC 迁移支持)。
Aug 5, 2026
媒体
The Decoder: Cursor in Slack improvements
The Decoder 2026-07-17 tech-press 报道 Cursor in Slack 改进(plan-upfront、multi-repo environments、cross-channel workflows)。
Aug 5, 2026
社区
Hugging Face blog: Security incident disclosure — July 2026
Hugging Face 7 月 16 日博客披露一起由「autonomous AI agent 系统」端到端驱动、对 HF 部分生产基础设施的入侵;攻击者利用 dataset processing 的代码执行路径(远程代码 dataset loader + dataset configuration 的 template-injection)在处理 worker 上跑代码,升级到 node 级访问,窃取云与集群凭证,并在周末横向移动到多个内部集群;整个攻击由一个 autonomous agent framework 执行,在 swarm of short-lived sandboxes 中跑了数万次单个动作。披露的影响:部分内部 datasets 未授权访问,若干 HF 服务凭证被攻陷,公开面向用户的 models / datasets / Spaces 无被篡改证据,软件供应链(container images / published packages)验证 clean。HF 使用 zai-org/GLM-5.2 在自家基础设施上做 LLM-based triage over security telemetry,分析超过 17,000 条事件记录,原因是商业 API frontier model 因安全护栏把响应者也当成攻击者拦截。社区建议:预防性 rotate access tokens 与审查近期账号活动;security@huggingface.co 接报。
Aug 5, 2026
论文
arXiv 2607.12395: Ring-Zero — Scaling Zero RL to a Trillion Parameters for Emergent Reasoning
arXiv 2607.12395(在 HF daily papers 上 7/16 索引)提出 Ring-2.5-1T-Zero —— 一个用 zero RL 在预训练 base 模型上直接训练、不需要人工标注 CoT 数据、依赖 RLVR 的模型;架构使用 Ling-2.5-1T-Base(1T MoE / 63B 激活)与 Ling-2.5-flash-Base(104B MoE / 7.4B 激活)两个 base 模型,四阶段训练 pipeline(一阶段 RL 用 clipped importance sampling policy gradient + token-level loss 诱发推理 / 自蒸馏用最短正确 rollout + 自反思压缩冗长 CoT / 二阶段 RL 用 sample-level loss normalization 做稳定优化 / 三阶段 RL 引入 4k / 16k / 64k 三层 token budget tier);基础设施用 320x H200 GPU + Megatron + SGLang + Areal 编排,BF16 body + FP32 attention softmax / LM head 混合精度,context parallelism 针对 hybrid MLA + Lightning Attention 架构调优。Ring-2.5-1T-Zero(Second Stage RL)在 AIME 2024 上达 94.1%,在 AIME 2025-2026 / HMMT 2025-2026 / IMOAnswerBench 上有竞争力;自发涌现 5 种认知行为(anthropomorphism / structured formatting / self-verification / parallel reasoning / context anxiety);CoT trace 在 comprehensibility / reproducibility(Qwen-32B 自蒸馏比 DeepSeek-R1 高 5.8 分)/ efficiency(用不到一半 token)上优于 baseline。
Aug 5, 2026
社区
vLLM blog: Keeping vLLM Production Quality
vLLM 7 月 16 日博客《Keeping vLLM Production Quality》记录每个 PR 必经的三层关卡:Layer 1 CI(每个 PR 跑轻量 GitHub Actions,Buildkite 上 37 个测试组 / 266 个任务按 diff 动态选取,共享多阶段容器镜像,pip-compile lock 文件,58 个 runner 队列跨 AWS / Crusoe / LambdaLabs / Nebius / NVIDIA / Roblox / RunPod,MIG-slicing 与 autoscale-from-zero 每机器 runner,自定义仪表盘 ci.vllm.ai,夜间 CI-analyzer bot 每天约 1.5 个 auto-revert PR / 约 70% 正确诊断);Layer 2 性能与精度(github.com/vllm-project/perf-eval 夜间 pipeline 跨 H200 / B200 / MI300X / MI355X 跑 17 个 model-hardware 配方:DeepSeek V4 Pro/Flash、gpt-oss、Kimi K2.5、MiniMax M2.5/M3、Qwen3.5、GLM 5.1、Gemma 4、Nemotron 3 Super,度量 TTFT / TPOT / vllm-bench,精度 lm-eval GSM8K / GPQA / AIME,函数调用 BFCL);Layer 3 release 自 2025 年 11 月起的双周节奏(每隔一周的周一 release manager 从 main 切 releases/vX.Y.Z,周一至周三 cherry-pick 窗口,只有三层全通过才发布候选),每个 release 发 7 个 Python wheels + 11 个 Docker images,发布前每个都做 smoke-test。
Aug 5, 2026
官方
Cohere blog: Cohere and the University of Toronto partner to advance responsible AI adoption at scale
Cohere 7 月 16 日博客宣布与多伦多大学的多年合作 —— 把 Cohere 的企业 AI 技术集成到 U of T 即将上线的全校 AI 平台,支持教学、研究、学生服务、行政、运营五大场景的负责任 AI 落地;Cohere 的 North 平台作为 U of T AI 平台内的 orchestration layer,帮用户管理复杂任务、安全访问各大学系统中的可信信息,支持 faculty / librarian / staff / student,同时把敏感数据留在大学的控制之下;Cohere 的技术也将驱动 U of T 的「AI Kitchen」—— 一个通过审核过的应用、合适的数据访问与隐私优先框架来探索和评估 AI 工具的安全环境;这次合作是 Cohere 创始人的回家 —— 2019 年他们以 U of T 学生身份创立公司(Aidan Gomez、Nick Frosst、Ivan Zhang);文章未披露具体资金数字。
Aug 5, 2026
社区
Hugging Face blog: Newer Models, Same Advantage — Dharma-AI
Hugging Face 7 月 16 日博客《Newer Models, Same Advantage》(Dharma-AI)在巴西葡萄牙语 OCR 任务上对比 Mistral OCR4(0.798)、Unlimited-OCR(0.7587)与 DharmaOCR(0.925);两个更新的通用 / 多语言模型虽然更新、资源更充足,但分数都明显低于 DharmaOCR;文章把 DharmaOCR 的持续优势归因于结构特化 —— (1) 领域集中:全部参数都用在巴西葡萄牙语词汇、形态、拼写模式上,而非分散在多种语言;(2) 两阶段训练:监督微调建立领域能力,再用 Direct Preference Optimization 在完整输出连贯性(而非逐 token 预测)上做训练,降低视觉歧义下的文本崩坏;文章主张这种优势是结构性而非临时性的 —— 通用模型把资源分散在多领域,专家模型用有限资源在本领域拿到更多;Mistral OCR4 把「Chico Buarque」读成「Chico Barque」,Unlimited-OCR 输出「a dose de chico bique」这种不连贯文本。
Aug 5, 2026
论文
arXiv 2607.13285: Harness Handbook — Making Evolving Agent Harnesses Readable, Navigable, and Editable
arXiv 2607.13285(在 HF daily papers 上 7/16 索引)是腾讯混元的一本 handbook / guide,讲如何设计更清晰、可导航、可编辑的 agent harness 架构。具体作者名单、架构模式目录、评估方法与参考实现均未提取。
Aug 5, 2026
论文
arXiv 2607.12747: Tracing Agentic Failure from the Flow of Success
arXiv 2607.12747(在 HF daily papers 上 7/16 索引)是一篇 4 作者论文,从成功的轨迹里反推 agentic 失败模式 —— 即在同一条轨迹曾经成功的环节定位失败。具体作者名单、失败分类法、评估方法与所用 agent / benchmark 均未提取。
Aug 5, 2026
论文
arXiv 2607.12625: KnowAct-GUIClaw — Know Deeply, Act Perfectly, Personal GUI Assistant with Self-Evolving Memory and Skill
arXiv 2607.12625(在 HF daily papers 上 7/16 索引)提出了 KnowAct-GUIClaw —— Lychee Team 的 GUI assistant,带 self-evolving memory 和 skill 模块。具体作者名单、Android / Web / Desktop GUI agent benchmark 上的具体分数、memory 模块架构、self-evolution 机制均未提取。
Aug 5, 2026
社区
Hugging Face: Security incident cluster thread
7 月 16 日 HF 安全事件三件套同日上:Hugging Face Security Incident Disclosure 主文(autonomous AI agent 端到端攻击 + GLM-5.2 取证)+ The Decoder tech-press 报道(autonomous agent 攻击向量、GLM-5.2 取证、合作伙伴 / 客户影响评估)+ Anatomy of a Frontier Lab Agent Intrusion 伴随深度技术时间线(HF blog 索引标注 security / Hot / 423 reactions)。三件套拼成「业界首例公开披露的 autonomous AI agent 攻击」完整叙事。
Aug 5, 2026
仓库
vLLM perf-eval GitHub repo (companion)
vLLM 7 月 16 日 perf-eval 仪表盘覆盖 17 个 model-hardware 配方,跨 H200 / B200 / MI300X / MI355X(DeepSeek V4 Pro/Flash、gpt-oss、Kimi K2.5、MiniMax M2.5/M3、Qwen3.5、GLM 5.1、Gemma 4、Nemotron 3 Super),度量 TTFT / TPOT / vllm-bench,精度 lm-eval GSM8K / GPQA / AIME,函数调用 BFCL;结果展示在 ci.vllm.ai/perf 与 ci.vllm.ai/eval。具体当前排行榜快照与每个 recipe 随时间的 TTFT / TPOT / 精度 delta 均未提取。
Aug 5, 2026
社区
Hugging Face: Welcome Inkling by Thinking Machines
Hugging Face 7 月 15 日博客宣布 Thinking Machines 的 Inkling —— 一个 ~1T 参数的开源多模态模型,原生接受图像、文本、音频输入,在 45T token 上训练。架构:decoder-only 多模态 MoE,975B 总参 / 41B 激活,256 个 expert + top-6 + 2 shared expert,1M 上下文窗口,相对注意力(无 RoPE),5:1 sliding-window 与 global 混合注意力,hidden state 上的短 1D 卷积(SConv),视觉用 hierarchical MLP patchifier,音频用 mel-spectrogram 离散化;变体为 Inkling BF16(2 TB VRAM)与 Inkling NVFP4(600 GB),Inkling-Small BF16 600 GB(276B / 12B 激活)/ NVFP4 180 GB;自报分 Inkling vs Inkling-Small:HLE 文本 29.7% / 31.6%、HLE 带工具 46.0% / 47.8%、AIME 2026 97.1% / 95.5%、GPQA Diamond 87.2% / 89.5%、SWE-Bench Verified 77.6% / 80.2%、SWE-Bench Pro 54.3% / 55.9%、Terminal Bench 2.1 63.8 / 64.69、MCP Atlas 74.1% / 79.2%、MMMU Pro 73.3% / 74.0%、VoiceBench 91.4% / 90.1%;协议未公开。
Aug 5, 2026
社区
vLLM blog: TML Inkling on vLLM — Day-0 Support
vLLM 7 月 15 日博客记录 Inkling day-0 支持 —— NVFP4 + BF16 变体完整功能对等,带 8 个 MTP head(每次 forward step 最多 9 个 token),原生支持 1M token;在 4x NVIDIA GB200 上 380 tok/s/user(MTP8,平均 acceptance length 4.5)、140 tok/s/user(无 MTP);按长度桶的精度数据 2K-221K 99.09% (436/440)、294K-513K 95.68% (421/440)、586K-805K 81.36% (358/440)。
Aug 5, 2026
社区
LMSYS blog: Inkling Day-0 Support in SGLang
LMSYS 7 月 15 日博客宣布 SGLang 同步 day-0 支持 Inkling —— 975B 多模态 MoE、1M 上下文,在 Blackwell 上达到 71.7k tok/s 输入吞吐。SGLang Cookbook 同步给出 Inkling 在 Blackwell TP4/TP8 / H200 / AMD MI350X / MI355X 的部署菜谱。
Aug 5, 2026
社区
Hugging Face: Real World VoiceEQ — Measuring the human quality of voice AI
Hugging Face 7 月 15 日博客宣布 Real World VoiceEQ —— 由 Hume + HF 共同构建的语音 AI 质量基准,基于 100 万+ 跨人群、口音、声学环境的人工评分,覆盖 40+ 专有与开源语音模型、15+ 评测维度、60+ 指标(ASR / TTS / S2S / 语音理解);当前数据集 78.5 万条 TTS 评分、4.8 万条 STS 评分;每次评测在 Hume 的 Kairos 平台运行;引用噪声环境下语音转写 WER 约为音乐环境下的 4 倍;公开 leaderboard 在 huggingface.co/spaces/HumeAI/rw-voice-eq,技术报告在 arXiv 2607.14846。
Aug 5, 2026
官方
Cohere blog: The total cost of AI ownership
Cohere 7 月 15 日博客《The total cost of AI ownership》把 AI 成本框定为「renting vs owning」跨数据中心、芯片、模型三层;引用 Gartner 2026 全球 AI 支出 2.52 万亿美元(YoY +44%)、IDC/DataRobot 96% 生成式 AI 与 92% agentic AI 部署面临高于预期成本、McKinsey ~1/3 组织把 AI 全企业规模化(5-6% 报告显著财务影响)、Mavvrik/Benchmarkit 80% 公司 AI 预测偏离 >25%、Uber 10% 已承诺代码由自主 agent 构建 12 个月预算 4 个月花光;NVIDIA Blackwell vs Hopper ~50x 每兆瓦 token / ~35x 每 token 成本;Lenovo 2026 摊销每百万 token 在自有 H100 ~$0.11、云实例 ~$0.89、前沿 API ~$2.00(对云 8x / 对 API 18x);8-GPU 服务器相对按需云在不到 4 个月内回本;收支平衡约在每天 4 小时使用率;SemiAnalysis InferenceX 在 GB300 上 $0.123 每百万 token。
Aug 5, 2026
社区
Hugging Face: Model Routing Is Simple. Until It Isn't. — IBM Research
Hugging Face 7 月 15 日博客文章,由 IBM Research 发表,题为《Model Routing Is Simple. Until It Isn't.》,讨论为什么随着系统规模增长,模型路由(把查询导向不同 AI 模型)会变得复杂,与初始的简单印象相违。具体框架名、决策图、代码样例、基准方法学在已捕获摘要里未提取。
Aug 5, 2026
社区
Hugging Face: Inkling day-0 cluster thread
7 月 15 日 Inkling 一日三发:Thinking Machines 在 HF 上发 ~1T 多模态开源 MoE、vLLM blog 发 day-0 支持(4x GB200 上 380 tok/s/user MTP8 + 1M 上下文)、LMSYS blog 发 SGLang day-0 支持(Blackwell 上 71.7k tok/s 输入吞吐);三方合起来把「1T 开源 + 一日就绪到推理引擎」做成了 7 月开源生态的最强一日之一。
Aug 5, 2026
论文
arXiv 2607.12463: Function-Aware Fill-in-the-Middle as Mid-Training for Coding Agent Foundation Models
arXiv 2607.12463(在 Hugging Face daily papers 上 2026-07-15 索引)提出了一种函数感知的 fill-in-the-middle 方法,用作 coding agent 基础模型的中训练(mid-training)。具体作者名单、HumanEval / SWE-Bench / LiveCodeBench 等基准的精确提升、训练算力开销、每步训练方法学在已捕获摘要里未提取。
Aug 5, 2026
论文
arXiv 2607.11562: MonkeyOCRv2 — A Visual-Text Foundation Model for Document AI
arXiv 2607.11562(在 Hugging Face daily papers 上 2026-07-15 索引)提出了 MonkeyOCRv2 —— 一个面向文档 AI 任务的视觉-文本基础模型。具体作者名单、DocVQA / InfoVQA / ChartQA / OmniDocBench 等基准的精确数字、模型规模、训练数据组成与协议在已捕获摘要里未提取。
Aug 5, 2026
官方
Anthropic newsroom: Claude for Teachers
Anthropic 7 月 14 日(7 月 21 日更新)新闻发布 Claude for Teachers —— 为美国已认证 K-12 教师免费提供 Claude 高级功能、教学 skill 库、Learning Commons 接入(覆盖 50 个州学术标准对齐);能力含差异化教学支持、Claude Code + Cowork 自主任务处理、数据分析工具(教师控制数据共享、不用于训练);平台集成 ASSISTments / Brisk Teaching / Canva Education / Coteach / Diffit / Eedi / MagicSchool / Snorkl / TeachFX,加 OpenSciEd + Illustrative Mathematics 课程资源;对已认证 K-12 教师免费,2027-06-30 前注册得一整年访问;18+ 政策、FERPA 通过 K-12 Data Processing Addendum 合规,Anthropic 正在与美国教师联合会(AFT)合作制定 Gold Standard。
Aug 5, 2026
社区
vLLM blog: vLLM x TileRT — Specialized Decode for Latency-Critical Serving
vLLM 7 月 14 日博客记录 vLLM × TileRT 集成 —— 通过 vLLM V1 公开 connector 接口把 prefill 留在 vLLM、decode 交给 TileRT;TileRT 0.1.5 随此集成发布;两池通过 MultiConnector 在单一 prefill 实例后共存,routing/claim filter 只把标 latency-critical 的流量导向 TileRT;评测在 8x NVIDIA B200 跑 GLM-5.1-FP8、输入 1K-192K / 输出 1K,MTP 平均 acceptance length 3.2(峰值 4.0);每 TileRT decode 节点同一时刻一个 in-flight 请求;模型覆盖 GLM-5/5.1 + DeepSeek-V3.2。
Aug 5, 2026
仓库
SGLang release v0.5.15.post1
SGLang 7 月 14 日发布 v0.5.15.post1 —— 针对 GLM-5.2 的稳定性补丁,修复非 CUDA / HIP 设备上的 DSA 模型启动、flashinfer 依赖、NaN 输出,以及 PD-disaggregation + context-parallel 设置下的 GLM-5.2 IndexShare;搭配 v0.5.15(7 月 10 日发布,Blackwell 上 GLM-5.2 NVFP4 生产调优、默认启用 Spec V2 +11% e2e TPS、IndexShare MTP、TopK V2、indexer-prologue fusion、新增 Hunyuan 3 / HRM-Text / Qwen3.6 NVFP4)。