A reported, fixed, and bountied security research project by Hacktron AI exposes a structural gap that most organisations carry without knowing it: bug-bounty scope and identity federation are different boundaries, and they do not have to overlap.
What Happened
To be explicit at the outset: this is responsible-disclosure security research. It was reported to the affected parties, fixed, and paid a bounty. It is not an attack, not a breach, and not a live incident. Internal OpenAI code repositories were not accessed, and no sensitive information was taken from any account. The Wall Street Journal first reported the research, subsequently amplified by the Financial Times; Hacktron AI published its own detailed account at hacktron.ai.
Three researchers at Hacktron AI spent two months chaining a sequence of known components into a single demonstrable path. The published chain ran from a heap buffer overflow in libheif, through a missing security backport in Debian 12, ImageMagick’s use of libheif for HEIC image conversion, the Discourse image upload pipeline, OpenAI’s community forum at community.openai.com, and an SSO misconfiguration in OpenAI’s identity infrastructure — reaching connected integrations including GitHub, Slack, and email access.
What the researchers demonstrated was narrow and they are explicit about its limits. They obtained an OpenAI employee’s Codex credentials and, in their words, “sent a prompt to this employee’s Codex account to open a PR for us in OpenAI’s internal monorepo” — without accessing actual code content — then halted testing. No individual employee is named in this analysis. The research was responsible disclosure in practice as well as in label.
The key insight: A bug-bounty scope is a commercial boundary — it defines what an organisation will pay to hear about. An identity federation is a technical boundary — it defines what credentials reach. There is no rule requiring the two to coincide, and in this case they did not. That gap is the structural story, not the vulnerability chain itself.

The Structural Read
There are four things worth separating cleanly here, because the noise around security research tends to collapse them.
First: the identity boundary versus the scope boundary. OpenAI’s own statement is the clearest evidence of the gap. The company said plainly that “testing against the Discourse-hosted community.openai.com was explicitly excluded from our bug bounty program. The award recognizes the OpenAI-side finding, not the actions against Discourse.” A community forum is not a production system, and that exclusion is a legitimate commercial decision — this analysis does not criticise it, call it a mistake, or suggest how any programme ought to be scoped. The structural point is simply that the scope said what OpenAI would pay to hear about; the identity federation said what credentials reach. The two described different perimeters. The chain moved between them.
Second: what was not done belongs alongside what was. The researchers opened a pull request through an employee’s Codex account and then stopped. Internal repositories were not accessed. No sensitive information was taken. No individual was harmed. Stating what did not happen is not a hedge — it is the accurate description of a responsible-disclosure project, and speculating about what might otherwise have been reachable would be irresponsible here.
Third: the response timeline. OpenAI confirmed a fix roughly fourteen hours after submission. Discourse moved from report to published advisory in three days. These figures measure different things — a fix confirmation on one side, a published advisory on the other — and are not a race. Neither is described here as fast or slow, and neither party is ranked.
Fourth: the model comparison, which requires the most careful labelling. The researchers used Claude Opus 4.8 for initial exploit development the researchers used Claude Opus 4.8 and, by their account, it “struggled across several sessions to produce a working exploit.” Claude Opus 5, released on 24 July 2026 partway through the project, by their account produced a working ARM64 exploit within three hours before porting it to x86-64. The task was held constant and the model changed underneath it, on real work, mid-project — which is exactly the comparison a benchmark tries to construct artificially.
Editorial label — read this before the comparison
This is one team, one task, self-reported, with no control condition and no repetition. It is an anecdote, not a measurement. Nothing here claims any capability level, score, or generational factor; nothing says any model is better at security work; nothing generalises beyond this single project. A same-task, same-team, before-and-after observation across two model generations is a rarer form of evidence than any leaderboard produces — but rarer does not mean representative.
Business Engineer — Marginal Cost Shift
When the binding constraint stops being budget
The researchers’ own account states: “The whole HEIF Heist research project…took two-months, cost less than $3,000 in tokens in total, and was conducted by three researchers.” That figure is tokens only. It excludes salaries, tooling, infrastructure, and everything else three people working for two months actually cost — and it should not be read as the total cost of the project. Held loosely, the structural observation is this: when the marginal cost of attempting a hard technical task falls far enough, the binding constraint stops being budget and becomes judgement about what is worth attempting at all. Nothing here predicts any increase in attacks, claims offensive research is now cheap in general, or extrapolates this figure to any other project.
Three Implications
IMPLICATION 1 — Scope Maps and Federation Maps Need to Be Read Together
Every organisation that runs a bug-bounty programme has a scope document and an identity architecture. This research makes visible that those two artefacts can describe different perimeters — and that the gap between them is a structural condition, not a negligence finding. The practical question for any security or engineering team is whether the two maps have ever been placed side by side. This is not a criticism of OpenAI’s programme design; it is a framework question that applies to any organisation running federated identity alongside a scoped disclosure programme.
IMPLICATION 2 — The Anecdotal Model Comparison Is Still Worth Tracking, With the Label On
The before-and-after observation — same team, same task, model changed mid-project — is the kind of evidence that leaderboards structurally cannot produce, because leaderboards hold the model constant and vary the task. That makes this anecdote interesting without making it generalisable. Organisations evaluating AI-assisted engineering workflows have more reason to run their own equivalent experiments on their own tasks than to read capability from any published benchmark. The label belongs on every time this comparison is cited: one team, one task, self-reported, no control, no repetition.
IMPLICATION 3 — The Constraint Shift From Budget to Judgement Is the Durable Strategic Signal
If the token cost of attempting a hard technical task falls to a level that no longer functions as a filter, the question an organisation — or a research team — asks changes from “can we afford to try this?” to “is this worth attempting?” That shift in the binding constraint has implications for how security teams prioritise defensive research, how platforms think about their own attack surface, and how disclosure programmes are resourced. No increase in attacks is predicted here; no general claim about the cost of offensive research is made. The constraint shift is a structural observation, not a forecast.
91,000+ executives read Business Engineer for the AI strategy frameworks cited by ChatGPT, Claude, and Perplexity.
This describes responsible-disclosure security research that was reported, fixed and paid a bounty. It is not an attack, not a breach and not a live incident. Internal OpenAI code repositories were not accessed, no sensitive information was taken from any account, and the researchers halted testing after the proof-of-concept. No technical steps, code, commands, payloads or reproduction detail appear above; the components are named only as the published chain names them. Nothing here speculates about what might otherwise have been reached, describes any hypothetical worst case, or quantifies any potential impact, and no individual employee is named or described. The comparison between two model generations is the researchers’ own account of one task in one project, self-reported, with no control condition and no repetition. It is an anecdote rather than a measurement. Nothing here claims any capability level, score or generational factor, says any model is better at security work, or generalises beyond this project. OpenAI stated that testing against the Discourse-hosted community forum was explicitly excluded from its bug bounty programme and that the award recognises the OpenAI-side finding. Nothing here criticises that scope decision, calls the exclusion a mistake, or suggests how any programme should be scoped. The two response times cited measure different things — a fix confirmation and a published advisory — and are not comparable as a ranking. Neither is described as fast or slow, and neither party is praised or criticised. The cost figure is a token cost only. It excludes salaries, tooling, infrastructure and everything else, is not the cost of the project, and is not extrapolated to anything. Nothing here predicts any increase in attacks or claims that offensive research is cheap in general. OpenAI and Anthropic are private companies; no valuation, share-price or market-capitalisation claim is made. This is business analysis. It is not security advice and not investment advice, no view is expressed on any security, and no recommendation is made.
Sources: hacktron.ai · wsj.com · venturebeat.com








