Verified · Sep 19, 2026
Independently verifiedResearchers used Claude Opus 5 to reach OpenAI's internal repos — disclosed through bug-bounty channels, and per the team no internal code was accessed
4 sourcesThe researchers' own account (Hacktron blog, September 13, 2026): 'On July 25, 2026, we chained two critical vulnerabilities to compromise multiple OpenAI employees' ChatGPT accounts. With these accounts, we could then access internal OpenAI repositories, and potentially many other connectors.' And the receipt: 'To prove we had in fact gained the access we believed without allowing ourselves to learn any sensitive information, we used the employee's Codex to open a PR #1186742 in OpenAI's internal monorepo openai/openai .' The chain: a July 23 review found HEIC/HEIF uploads routed from Discourse through ImageMagick into libheif, where 'an heap buffer overflow' (verbatim) during HEIC decoding gave attacker-controlled input a path — a bug fixed upstream the previous year but never flagged as a vulnerability, so it 'received no CVE.' The model layer, the blog's words: 'On July 24, we used Opus 4.8 to develop a working ImageMagick/libheif code-execution exploit with ASLR disabled' — but making it reliable against ASLR-enabled defaults 'wasn't fruitful.' 'That evening, Anthropic released Claude Opus 5.' 'We started a new session, which first produced a working ARM64 exploit for a local Mac within 3 hours.' Both outlets carry the summary: 'Opus 4.8 struggled across several sessions to produce a working exploit with ASLR enabled. Within hours of Opus 5's release, we gave it the same problem and it succeeded.' (TechCrunch's quote of this drops the trailing 'with ASLR enabled' — the blog's fuller sentence is the canonical one.) TechCrunch (September 18, 7:00 AM PDT) adds the authorization frame: 'A three-person security team at startup Hacktron AI carried out the attack as part of an OpenAI bug-bounty program. Hacktron reported its findings to OpenAI, which gave the startup a $6,500 award.' 'OpenAI says it has resolved the issues Hacktron uncovered'. OpenAI's own scope comment, quoted in the blog's timeline: 'testing against the Discourse-hosted community.openai.com was explicitly excluded from our bug bounty program. The award recognizes the OpenAI-side finding, not the actions against Discourse.' Escalation root cause, the blog's emphasis: 'It is an OpenAI SSO issue that turned the forum compromise into access to ChatGPT and Codex.' Timeline: Bugcrowd submission July 25, OpenAI confirming the fix 'roughly 14 hours after the initial submission', Discourse reported the same day via HackerOne and fixed July 27, advisory July 28, bounty paid and resolved September 1 — and the blog's span sentence: 'The entire timeline from initial discovery to access to OpenAI repo access took place in less than 72 hours.' Anthropic's own page (July 24, 2026) supplies the safeguards line: its classifiers 'block "binary-based" vulnerability scanning ... penetration testing, and exploit generation', flagged requests fall back to Opus 4.8, and CVP members get 'a version of Opus 5 with fewer security restrictions' — while the blog notes 'Opus refused write exploit for remote instances' and that 'skilled human guidance remained important'. TechCrunch credits the story's mainstream break to 'The Wall Street Journal reported on Thursday evening' (September 17).
Why now
Three clocks converge on this weekend. The story's mainstream break was Thursday evening — 'The Wall Street Journal reported on Thursday evening', per TechCrunch — and TechCrunch's own write-up followed Friday morning at 7:00 AM PDT, so the audience is meeting the headline ('Claude hacked OpenAI') right now while the researchers' full account has been public since September 13 waiting to be read properly. Second, the debate the story feeds — AI collapsing the cost of elite exploitation — has two freshly quotable anchors: a security CEO's '$200 a month' line and a founder's 'Work that once took months can now take days.' Third, the card's differentiator has a deadline: the headline frame is already set, and the record underneath it — a chain disclosed through OpenAI's Bug Bounty Program on Bugcrowd and Discourse's HackerOne program, with the OpenAI-side finding carrying the $6,500 award and OpenAI's own scope comment, no data taken, both fixes dated in July — is what separates the explainer from the aggregation wave before it hardens into 'OpenAI got breached.'
Why it is worth publishing
Broadest tech-audience interest on the card, and the highest precision burden — which is the recommendation: this is the story where accuracy IS the content. The loaded verbs must stay in their lanes: 'hack into' and 'break into' are TechCrunch's frame of an authorized action; 'compromise' is the researchers' own verb; 'broke containment' belongs to a different, earlier incident (OpenAI's own agents in a cyber eval) that TechCrunch cites only as context. The model-capability claim belongs to the researchers ('we gave it the same problem and it succeeded'), while Anthropic's own page documents blocked exploit generation, Opus 4.8 fallbacks, and the CVP's fewer-restrictions access — no opened source says which channel the team used, so both 'the safeguards failed' and 'the model went rogue' are unsupportable. And the numbers all carry their welds: $6,500 as an award, 'less than 72 hours' per the blog's own endpoints, 'roughly 14 hours' for the fix confirmation, and the HEIF Heist '$3,000' belonging to a different, broader campaign.
Evidence basis
Four source records: Hacktron's blog (metadata datePublished September 13, 2026; read in full, raw HTML fetched, every quoted span grepped verbatim), TechCrunch (Aditya Mehta and Rebecca Bellan, Sept 18, 2026, 7:00 AM PDT; read in full, raw HTML fetched, same grep), and Anthropic's Opus 5 announcement (page-stated July 24, 2026; read in full, same grep) were all opened this run; the Wall Street Journal piece is an UNOPENED origin (paywalled; URL from TechCrunch's inline link), recorded under the double-relay rule with every fact from it riding TechCrunch's citation and no signal created for it. Weekday-date pairs calendar-verified: July 23 = Thursday, July 24 = Friday, July 25 = Saturday, July 26 = Sunday, July 27 = Monday, September 13 = Sunday, September 17 = Thursday, September 18 = Friday; today = Saturday, September 19, 2026. Numbers come in four welded clusters: the $6,500 award (both primary sources); the blog's spans ('less than 72 hours' initial discovery to repo access; 'roughly 14 hours' fix confirmation; 'within 3 hours' local Mac ARM64 exploit); Fredrikson's '$200 a month' quote; and the HEIF Heist campaign's 'less than $3,000' / two-months figures, which belong to the broader research, never to this operation.
“Three researchers say they used Claude Opus 5 to reach OpenAI's internal repos — disclosed through bug-bounty channels, they say, with no internal code accessed.”
Angle
Make it a 'read the security headline like an editor' explainer. Beat one: what actually happened in plain chain form — a crafted image, a library bug with no CVE, the forum server, an SSO flaw, employee accounts, a pull request as proof. Beat two: the authorization spine — paid bug bounty, $6,500 award, OpenAI's own scope comment, the team's no-data line quoted verbatim as the receipt. Beat three: the model layer with attribution welded on — Opus 4.8 struggled (per the researchers), Opus 5 succeeded within hours (per the researchers), and Anthropic's own page says exploit generation is blocked on its surfaces with CVP members getting fewer restrictions; nobody's opened source says which channel the team used. Beat four: the takeaways that age well — a year-old fix with no CVE stayed dangerous, the escalation was the SSO design not Discourse alone, and self-hosters are still being told to rebuild. Close on the economics quotes, attributed: '$200 a month' (Gray Swan's CEO, per TechCrunch) and 'months ... days' (Hacktron's founder, on X).
Format
Long-form explainer
Demo idea
A 'headline vs record' split-screen: the aggregation headline ('Claude hacked OpenAI') on the left; on the right, five record lines each with its date — bug bounty program (July 25 submission), $6,500 award (September 1, paid and marked resolved), 'without allowing ourselves to learn any sensitive information' (the blog, verbatim), OpenAI fix confirmed (~14 hours, July 25), Discourse fix (July 27) — footer: 'per the researchers' own account; OpenAI says resolved.'
Platform notes
This is cybersecurity content — say the authorization in the same sentence as the action, every time: paid bug bounty, reported, fixed, rewarded, testing stopped voluntarily. Never say OpenAI 'got breached' or that data was stolen: the record states no sensitive information was learned and no internal code was accessed, and 'steal' appears nowhere in TechCrunch's article. Loaded verbs stay in their lanes: 'hack into'/'break into' are TechCrunch's frame (label them); 'compromise' is the researchers' verb; 'broke containment' belongs to a different, earlier OpenAI-agents incident — importing it here fabricates a second breach. Give OpenAI's scope comment whenever the $6,500 comes up ('recognizes the OpenAI-side finding, not the actions against Discourse'). Attribute the Opus 4.8/Opus 5 account to the researchers, quote Anthropic's own safeguard lines when the safety question comes up, and say plainly that no opened source states which access channel the team used. Weld every number: $6,500 award, 'less than 72 hours' (blog's own endpoints), 'roughly 14 hours' (fix confirmation), 'within 3 hours' (local Mac exploit), '$200 a month' (Fredrikson's quote, per TechCrunch) — and never merge the HEIF Heist '$3,000' into this operation. Don't compute the 72 hours yourself from other dates; and don't write the two-months-ago exposure line in the present tense — both fixes landed in July, while the self-host rebuild warning is still standing.
Usable claims
- A two-vulnerability chain, disclosed through OpenAI's Bug Bounty Program on Bugcrowd and Discourse's HackerOne program, reached from OpenAI's community forum into the company's internal repositories, with OpenAI awarding $6,500 for the OpenAI-side finding — per the researchers' own account on the Hacktron blog (published September 13, 2026), TechCrunch's report (Aditya Mehta and Rebecca Bellan, 7:00 AM PDT, September 18, 2026), and Anthropic's Opus 5 announcement (July 24, 2026). Authorization and outcome, TechCrunch's words: 'A three-person security team at startup Hacktron AI carried out the attack as part of an OpenAI bug-bounty program.' 'Hacktron reported its findings to OpenAI, which gave the startup a $6,500 award.' 'OpenAI says it has resolved the issues Hacktron uncovered'. The researchers' account: 'OpenAI also paid us a $6,500 bounty.' OpenAI's own scope comment, as quoted in the Hacktron blog's timeline: 'To clarify the scope of that award: testing against the Discourse-hosted community.openai.com was explicitly excluded from our bug bounty program. The award recognizes the OpenAI-side finding, not the actions against Discourse.' Impact, the blog's words: 'On July 25, 2026, we chained two critical vulnerabilities to compromise multiple OpenAI employees' ChatGPT accounts. With these accounts, we could then access internal OpenAI repositories, and potentially many other connectors.' Proof without reading code: 'To prove we had in fact gained the access we believed without allowing ourselves to learn any sensitive information, we used the employee's Codex to open a PR #1186742 in OpenAI's internal monorepo openai/openai.' TechCrunch's parallel framing: 'The team managed to chain together two critical vulnerabilities to gain access to multiple OpenAI employee ChatGPT accounts, which gave them entry into the company's software.' Entry vector, the blog's words: 'On July 23, we started reviewing Discourse's image-upload pipeline, and we found that HEIC and HEIF files followed an unusual path. Discourse normally used FastImage for image checks, but because FastImage did not support HEIF, it passed those files to ImageMagick's magick command for conversion.' 'That exposed the underlying libheif parser directly to attacker-controlled files.' The bug: 'This allowed an heap buffer overflow leading to OOB R/W primitives during HEIC decoding.' Why it was still open: 'the vulnerable code had been changed upstream the previous year, but the commit was not documented as a security fix and received no CVE' — TechCrunch: 'But the fix was never formally flagged as a vulnerability, meaning it never got a CVE (common vulnerabilities and exposures) number'. TechCrunch's rendering of the effect: 'Buried inside libheif was a memory bug that exposed a path for an attacker to sneak in their own instructions. In this case, feeding the library a specially crafted image caused it to miscalculate where one image was positioned on top of another, which proved enough to hijack the server.' The model layer, the blog's words: 'On July 24, we used Opus 4.8 to develop a working ImageMagick/libheif code-execution exploit with ASLR disabled. We then launched several separate sessions to make it reliable against Discourse's default configuration with ASLR enabled, which wasn't fruitful.' 'That evening, Anthropic released Claude Opus 5.' 'We started a new session, which first produced a working ARM64 exploit for a local Mac within 3 hours.' — and 'By 6:00 a.m. on July 25, we had confirmed local RCE through an image upload.' The summary sentence both outlets carry: 'Opus 4.8 struggled across several sessions to produce a working exploit with ASLR enabled. Within hours of Opus 5's release, we gave it the same problem and it succeeded.' (TechCrunch quotes this passage without the trailing 'with ASLR enabled' — the blog's fuller sentence is the canonical one.) TechCrunch adds, attributed to the researchers: the Claude model they were using was 'a special version of Opus 4.8 made available for cybersecurity researchers' — that phrasing appears in TechCrunch's report, not in the blog's own text. Anthropic's own page, for the release fact and the safeguard line: 'Claude Opus 5 is available today.' (page-stated July 24, 2026); 'we've intentionally avoided training Opus 5 on cyber tasks'; and the classifiers 'allow Opus 5 to find vulnerabilities in source code, but block "binary-based" vulnerability scanning (a method more likely to be associated with malicious actors), penetration testing, and exploit generation.' The blog's own guardrail note: 'Opus refused write exploit for remote instances', and 'skilled human guidance remained important' — no opened source states which access channel or program tier the team used for the Opus 5 sessions. Escalation root cause, the blog's emphasis: 'We want to emphasize that the vulnerability to escalate is not Discourse-specific. It is an OpenAI SSO issue that turned the forum compromise into access to ChatGPT and Codex.' Timeline (all from the blog's timeline table, weekday-date pairs calendar-verified): July 23 (Thursday) review begins; July 24 (Friday) Opus 4.8 attempts, Opus 5 released that evening; July 25 (Saturday) 05:00–06:00 UTC initial RCE and admin access, 08:00–10:00 UTC Bugcrowd submission, 13:30–15:30 UTC employee-account proof-of-concept with testing ceased 'at approximately 15:30 UTC', 22:49:45 UTC OpenAI confirming the fix 'roughly 14 hours after the initial submission'; Discourse reported the same day via HackerOne, replied July 26 (Sunday), 'had a fix ready by Monday' July 27 — TechCrunch: 'which issued a fix on July 27' — advisory GHSA-vhm9-85gw-x335 published July 28; bounty paid and marked resolved September 1. The blog's span sentence: 'The entire timeline from initial discovery to access to OpenAI repo access took place in less than 72 hours.' The mainstream wave, TechCrunch's attribution: the story broke via 'The Wall Street Journal reported on Thursday evening' (September 17). The blog's standing warning: 'If you self-host Discourse, rebuild your installation now.'
Evidence pipeline
From the news
Breakdown
The heaviest editing job on today's card is keeping every loaded word attached to the event its source ties it to. The verb map: 'hack into' and 'break into' are TechCrunch's frame — its own second paragraph says the team 'carried out the attack as part of an OpenAI bug-bounty program'; 'compromise' is the researchers' own verb for what they did; 'hijack' belongs only to the libheif memory bug's effect on the Discourse server; 'took over' belongs only to the employee-account step; and 'broke containment' belongs to a different, earlier incident — OpenAI's own agents during a cybersecurity evaluation — that TechCrunch cites as context and that must never migrate onto this story. The award has fine print: OpenAI's own comment, quoted in the blog's timeline, says 'testing against the Discourse-hosted community.openai.com was explicitly excluded from our bug bounty program. The award recognizes the OpenAI-side finding, not the actions against Discourse' — the Discourse-side report went through HackerOne, so 'OpenAI paid for the whole thing' erases the record. The model layer has a three-way split: the success account is the researchers' own ('Within hours of Opus 5's release, we gave it the same problem and it succeeded'); TechCrunch's 'a special version of Opus 4.8 made available for cybersecurity researchers' is attributed to the researchers and appears nowhere in the blog's text; and Anthropic's own page documents the opposite-direction safeguards — classifiers that 'block ... penetration testing, and exploit generation', Opus 4.8 fallbacks, and CVP members' fewer-restrictions access — while the blog notes 'Opus refused write exploit for remote instances.' No opened source says which channel the team used; the honest card says so, and neither 'the safeguards failed' nor 'the model went rogue' survives the record. The numbers arrive pre-welded: $6,500 as an award; 'less than 72 hours' with the blog's own endpoints; 'roughly 14 hours' for OpenAI's fix confirmation; 'within 3 hours' for a local Mac exploit; '$200 a month' inside a CEO's quoted line; and the HEIF Heist '$3,000' belonging to a different, broader campaign. And the record's receipts run the other way from the headline: no sensitive information learned, no internal code accessed, testing stopped voluntarily, OpenAI fixed the same day (confirmed 'roughly 14 hours after the initial submission'), Discourse fixed July 27, bounty paid September 1. Editor's rule: authorization ships in the sentence with the action, verbs stay in their lanes, the fine print rides with the award, and the safeguard question is answered with 'no opened source says' — not with a side.
Sources
Risks
- Say the authorization in the same sentence as the action: a paid bug-bounty test, reported, fixed, and rewarded — with the team's own no-data line ('without allowing ourselves to learn any sensitive information') as your receipt. Quote TechCrunch's 'hack into' only as TechCrunch's headline frame. Give OpenAI its scope comment ('recognizes the OpenAI-side finding, not the actions against Discourse') whenever you mention the $6,500. Attribute the Opus 4.8-vs-Opus 5 account to the researchers, and when the safeguard question comes up, read Anthropic's own lines (blocked exploit generation, CVP, fallbacks) and say honestly that no opened source states which channel the team used. Keep every number welded to its qualifier, date both fixes (July 25 OpenAI, July 27 Discourse), and close with the live advice for self-hosters.
Demo ideas
- Headline-vs-record split screen: the aggregation headline on the left, five dated record lines on the right (bounty submission, award scope comment, no-data line, OpenAI fix ~14h, Discourse fix July 27)
- Timeline strip July 23 → September 1: review starts → ASLR wall → Opus 5 released that evening → ARM64 exploit in 3 hours → RCE and Bugcrowd submission July 25 → OpenAI fix confirmed in ~14 hours → Discourse fixed Monday July 27 → advisory July 28 → bounty paid and resolved September 1, each node one verbatim fragment with its source named