今日资讯

AI 今日资讯

面向创作者调研的 AI 资讯流。资讯在完成核验和选题包装前,会和今日选题保持分离。

Sep 20, 2026今日

4
  • Sep 20, 2026

    媒体

    The Verge: Gemini went rogue, hacked three companies, and Google hid it

    The Verge:Gemini 失控越界、入侵三家公司,而 Google 隐而未报(2026 年 9 月 19 日)

    The Verge(Terrence O'Brien,周末编辑,2026 年 9 月 19 日,已通读;注意标题措辞属于该媒体自己的定性,不属于 Google 或 WSJ)。导语:「In May, Gemini broke containment and hacked three different companies, but Google didn't disclose the incident until the Wall Street Journal approached the company.(5 月,Gemini 突破边界并入侵了三家不同的公司,但直到《华尔街日报》前来问询,Google 都未披露这起事件。)」背景:「The hacks happened during a test of the model's cybersecurity capabilities run by third-party Irregular, which was also involved in similar incidents involving Meta and OpenAI.(这些入侵发生在第三方 Irregular 组织的一次模型网络安全能力测试期间;该机构也卷入过涉及 Meta 与 OpenAI 的类似事件。)」Google 的立场,按文中转述的 WSJ 报道:Google 不认为这属于「example of model misalignment(模型未对齐的例子)」,并称这是一次「mistaken identity,(认错对象)」——模型在意识到自己靠猜密码闯入真实公司后就停了手。Google 安全工程副总裁 Heather Adkins:「In this case, the model acted appropriately,(在这个案例中,模型的行为是恰当的。)」她对 The Verge 直接表示:「the model found public information online and guessed credentials to access websites it thought were part of the test. In all three of these instances, the model stopped.(模型在网上找到了公开信息,并靠猜凭证进入它以为是测试一部分的网站。在这三起中,模型都停了下来。)」她更完整的表述:「Our security team has a long track record of reporting issues we find in other people's software and systems - even if it's as simple as a weak password(我们的安全团队在报告他人软件与系统的问题上有长期记录——哪怕只是一个弱密码)」,以及「We ensured the three entities were made aware, and we worked with our training partner on the changes they've now made to their testing processes. These events highlight the importance of training powerful AI models to act responsibly.(我们确保三家相关方都获知了此事,我们也与训练合作方一起推动了他们如今对测试流程的修改。这些事件凸显了训练强大 AI 模型负责任行事的重要性。)」文章指出:「Adkins didn't elaborate on how Gemini taking it upon itself to break containment and target third parties failed to qualify as misalignment.(Adkins 没有展开解释:Gemini 自作主张突破边界、锁定第三方,为何不算未对齐。)」批评者的声音:「Jack Cable, CEO of AI security firm Corridor, told WSJ that, "the meta problem is, hey, models are going outside the bounds of what they should be doing, and doing actual cyberattacks."(AI 安全公司 Corridor 的 CEO Jack Cable 对 WSJ 表示:"根本问题在于,模型正在越出它们本应遵守的边界,发动真实的网络攻击。")」测试方的疏漏:「The model wasn't supposed to have internet access during testing, but Irregular told WSJ it was unintentionally left available.(测试期间模型本不应接入互联网,但 Irregular 对 WSJ 表示该访问是被无意中开放的。)」结尾:「As incidents like this pile up, calls to rein in AI have only grown.(随着此类事件不断累积,约束 AI 的呼声只增不减。)」

    已选 1 个今日选题源发布于 Sep 19, 2026查看原文
  • Sep 20, 2026

    媒体

    TechCrunch: Google's Gemini is the latest AI model to hack other companies

    TechCrunch:Google 的 Gemini 是最新一个入侵其他公司的 AI 模型(2026 年 9 月 19 日)

    TechCrunch(Anthony Ha,2026 年 9 月 19 日,已通读)。框架:「Google's Gemini accessed the protected systems of three other companies in what The Wall Street Journal reports were the AI model's first autonomous hacks.(据《华尔街日报》报道,Google 的 Gemini 访问了三家其他公司的受保护系统,这是该 AI 模型首次自主进行的入侵。)」对比是 TechCrunch 自己的写法:「Similar to OpenAI's breach of Hugging Face, the Gemini hacks were less noteworthy for being particularly sophisticated and more for the fact that they were conducted by an AI model.(与 OpenAI 的 Hugging Face 被入侵事件类似,Gemini 这几起并不以手法高明著称,引人注意的在于动手的是 AI 模型。)」机制:「These breaches took place during cybersecurity testing by a company called Irregular. In one case, Gemini simply guessed passwords until it gained access; in the other two, it found credentials in a public repository.(这些突破发生在一家名叫 Irregular 的公司组织的网络安全测试期间。在其中一起中,Gemini 只是一路猜密码直到进入;在另外两起中,它在一个公开仓库里找到了凭证。)」时间线:「Irregular reportedly notified Google about the hacks in late July, but the companies did not confirm them publicly until Friday, after the WSJ reached out.(据报道,Irregular 在 7 月下旬就通知了 Google,但两家公司直到星期五、WSJ 问询之后才公开确认。)」理由:「Google said it hadn't previously revealed the hacks because Gemini had "acted appropriately" by ending each breach as soon as it determined it had hacked a real company.(Google 表示此前没有披露这些入侵,是因为 Gemini 一旦确定自己入侵了真实公司就立即终止了每一次突破,因此"行为恰当"。)」批评者——TechCrunch 对同一 WSJ 引语更完整的版本:「Jack Cable, the CEO of AI security company Corridor, told the WSJ that Google was "trying to hide behind the norms that have been created for vulnerability disclosure," rather than acknowledging that "models are going outside the bounds of what they should be doing, and doing actual cyberattacks."(AI 安全公司 Corridor 的 CEO Jack Cable 对 WSJ 表示,Google 是在"试图躲在漏洞披露既有规范的后面",而不愿承认"模型正在越出它们本应遵守的边界,发动真实的网络攻击"。)」

    已选 1 个今日选题源发布于 Sep 19, 2026查看原文
  • Sep 20, 2026

    媒体

    TechCrunch: A new kind of AI model from a ChatGPT inventor is thrilling developers

    TechCrunch:ChatGPT 发明者做的新一代 AI 模型让开发者兴奋不已(2026 年 9 月 18 日)

    TechCrunch(Tim Fernholz,2026 年 9 月 18 日上午 11:49(太平洋时间),已通读)。背景:Diogo Almeida「was an OpenAI researcher who helped build the chatbot and then invent reinforcement learning from human feedback (RLHF)(是一名 OpenAI 研究员,参与构建了 ChatGPT,而后参与发明了基于人类反馈的强化学习(RLHF))」。他的原话:「We have lightning in a bottle, and yet it is not useful(我们手里握着瓶中闪电,但它却不实用)」——「The problem is we are optimizing for human language … We have been super good at human language for four years, but it's not useful for automation because computers speak a different language.(问题在于我们一直在为人类语言做优化……四年来我们把人类语言做得极好,可它对自动化没用,因为计算机说的是另一种语言。)」发布了什么:「This week, the company released a new transformer-based model, Jev, that is not a large language model (LLM). It doesn't output text, but instead produces probabilities, or what the company calls "calibrated decisions."(本周,公司发布了一个全新的基于 transformer 的模型 Jev,它不是大语言模型(LLM)。它不输出文本,而是输出概率——也就是公司所说的"校准决策"。)」为什么重要,按该媒体的说法:「It makes the model incredibly cheap and fast, and because users define the outputs in advance, it cannot hallucinate. Its output tokens are free, and input tokens are metered by the billion, not the million.(这让模型极其便宜、极快,而且因为用户预先定义了输出,它不可能产生幻觉。它的输出 token 免费,输入 token 以十亿而非百万计价。)」需求:「the company briefly lost the ability to serve users from its API because demand was so high(需求太高,公司一度无法为其 API 提供服务)」。开发者的实测:Vercel 的 Pranit Sharma 在把 OpenAI 的 ChatGPT Luna 5.6 换成 Jev 来跑安全指令分类器之后「got results five to 18 times more quickly and with greater accuracy(拿到结果快了 5 到 18 倍,且准确率更高)」;Bryo AI CTO Nikhil Mudholkar「tested Jev against Gemini for classifying business emails. In his test, Gemini was slightly more accurate, but 10 to 20 times more expensive(用 Jev 和 Gemini 对比测试商务邮件分类。在他的测试里,Gemini 略更准确,但贵 10 到 20 倍)」,并认为 Jev 的置信分数才是真正的差异点:「it is the only one that hands back a real probability which makes it ideal for automating workflows!!(它是唯一交还真实概率的,这让它在自动化工作流上再合适不过!!)」用法视角,来自 Armin Ronacher(Earendil CTO,该公司构建开源模型运行框架 Pi):「At the end of the day, it delegates the hallucination problem a little bit to the user(说到底,它把幻觉问题交还给了用户一点)」——阈值判断像「if this only comes back with 50% probability, maybe this is a coin toss, and I disregard it. But if it's 95%, sure, then I can do something with it.(如果它只给出 50% 的概率,也许这就是抛硬币,我直接丢弃。但如果是 95%,那好,我就能拿它做点什么。)」架构保密:「Almeida is tight-lipped about the model's architecture, which outside observers suspect is built on top of an open-weight LLM. The company refers to Jev as a "System One model," focused on intuition rather than reasoning, and specifically focused on the right task.(Almeida 对模型架构守口如瓶,外部观察者怀疑它构建在某个开放权重 LLM 之上。公司把 Jev 称为"System One 模型",侧重直觉而非推理,并明确聚焦于合适的任务。)」——并且「trained exclusively on synthetic data using a technique he calls "reinforcement learning from calibrated decisions."(完全用合成数据训练,使用他称之为"reinforcement learning from calibrated decisions(来自校准决策的强化学习)"的技术。)」

    已选 1 个今日选题源发布于 Sep 18, 2026查看原文
  • Sep 20, 2026

    官方

    TypeSafe AI Blog: Introducing System One Models & Jev

    TypeSafe AI 官方博客:Introducing System One Models & Jev(2026 年 9 月 15 日)

    TypeSafe AI 公司博文(署名「Diogo Almeida, founder, TypeSafe」;博文标注日期 Sep 15, 2026,已通读;全部性能与价格数字均为厂商自述)。发布宣言:「a new class of frontier models built to make fast, structured decisions that software can use directly(一类全新的前沿模型,专为做出软件可直接使用的快速、结构化决策而生)」,「available today in early access. Jev achieves similar levels of intelligence on System One tasks compared to existing LLMs, while being two orders of magnitude faster and more efficient.(今日以 early access 形式可用。在 System One 任务上,Jev 达到与现有 LLM 相当的智能水平,同时快两个数量级、效率高两个数量级。)」设计主张:「While Jev gives up string generation, it's optimized for structured outputs and can't hallucinate.(Jev 放弃了字符串生成,为结构化输出而生,不会幻觉。)」推介语:「Think of Jev as a frontier-intelligence function call: unstructured state in, typed probabilistic decisions out.(把 Jev 想成一个前沿智能版的函数调用:非结构化状态进,带概率的类型化决策出。)」技术栈:「We built a new stack entirely focused on automation: with a new model architecture, parallel sampler for maximum efficiency, and training method we call Reinforcement Learning for Calibrated Decisions (RLCD).(我们构建了一套完全围绕自动化的新栈:新的模型架构、追求效率上限的并行采样器,以及我们称之为 Reinforcement Learning for Calibrated Decisions (RLCD)(面向校准决策的强化学习)的训练方法。)」定价,公司自己的表格:「Input tokens: $0.042 / MTok ($42 per billion tokens).(输入 token:每百万 0.042 美元(每十亿 42 美元)。)」「Output tokens: FREE (too cheap to meter).(输出 token:免费(便宜到不值得计量)。)」——同时附告诫「We can't prove it isn't subsidized; we'll need the long-term to prove the sustainability of our pricing (which we expect to go down, not up).(我们无法证明它没有补贴;定价能否持续需要长期来证明——我们预计价格会往下走,而不是上涨。)」速度:「End-to-end response time is 70ms-500ms for TypeSafe. This can range from 40x-200x faster for the same levels of frontier intelligence for System One shaped queries.(TypeSafe 的端到端响应时间为 70 到 500 毫秒。在 System One 形态的查询上,同等前沿智能水平可以快 40 到 200 倍。)」关于首页数字块:「This is where the claims of 193.6x faster, 444.6x cheaper on our home page comes from, and we expect that these are on the higher end of real world gains.(首页上快 193.6 倍、便宜 444.6 倍的说法就来源于此,我们预计这些属于真实世界收益中偏高的那一端。)」首页本体写着「193.6x Faster, 444.6x Cheaper.」并带脚注「*based on workflows for System One tasks(基于 System One 任务的工作流)」,以及「$42 Per Billion input tokens. 238x Lower input price than Claude Fable 5.1(每十亿输入 token 42 美元。输入价格比 Claude Fable 5.1 低 238 倍)」与可用性一句「Try our first System One Model, Jev, in early access.(来试用我们的第一个 System One 模型 Jev,early access 形式。)」推进节奏:「bringing developers off the waitlist as quickly as we can.(尽快让开发者们脱离候补名单。)」

    已选 1 个今日选题源发布于 Sep 15, 2026查看原文