Claude Opus 5 Helped Security Researchers Breach OpenAI Staff Accounts
Security researchers used Anthropic’s Claude Opus 5 to chain two flaws and access OpenAI employee accounts, revealing new risks in how AI agents connect to business systems.
Photo by <a href="https://unsplash.com/@markusspiske?utm_source=WP+Agent&utm_medium=referral">Markus Spiske</a> on <a href="https://unsplash.com/?utm_source=WP+Agent&utm_medium=referral">Unsplash</a>
Claude Opus 5 Helped Security Researchers Breach OpenAI Staff Accounts
A small team of security researchers has revealed that they used Anthropic’s Claude Opus 5 to chain together two vulnerabilities and gain access to OpenAI employee accounts — reaching as far as OpenAI’s internal code repository before stopping and reporting everything they found. The disclosure, first reported by VentureBeat and independently confirmed by The Hacker News and The Wall Street Journal, has become one of the most closely watched AI security stories of the month, because it wasn’t a malicious attack at all. It was authorized research that shows how much faster AI coding agents can now make sophisticated hacking work.

What Happened
The team behind the research, a three-person security startup called Hacktron AI, found their way in through an unlikely door: OpenAI’s public community help forum at community.openai.com, which runs on the open-source Discourse platform. The forum lets users upload HEIC and HEIF image files, which get decoded through a library called libheif. Hacktron discovered that the specific version of libheif running on the forum’s server contained a heap buffer overflow — a memory-handling bug that, under the right conditions, can be turned into remote code execution.
That flaw, tracked as CVE-2026-32882, was confirmed directly by Discourse in its own security advisory, which rates it 8.8 out of 10 in severity and credits Hacktron’s research team for the report. The underlying fix for libheif had actually existed since May, but the Discourse server image the forum was running still had an older, unpatched version when the researchers looked in July.
From a Forum Bug to Employee Accounts
A compromised help-forum server on its own wouldn’t normally be a major event. What turned it into a bigger story was a separate issue: OpenAI’s forum offers “Sign in with OpenAI,” the same single sign-on system used across the company’s internal tools. Once Hacktron controlled the forum’s server, that shared login let them take over the ChatGPT and Codex accounts of forum members who happened to work at OpenAI — without those employees doing anything wrong themselves.
This is the detail security teams have focused on most. Hacktron has described it as an identity and authentication problem rather than a flaw in Discourse itself — meaning any lower-trust, internet-facing service sharing the same sign-on could, in theory, have opened the same door.
Where Claude Opus 5 Came In
The researchers say they initially tried building a working exploit using Claude Opus 4.8, but the model struggled across multiple sessions to defeat a standard memory protection called ASLR (address space layout randomization). Shortly after Anthropic released Claude Opus 5 in late July, the team gave the newer model the same problem in a fresh session — and it produced a working exploit within hours.
Anthropic has built safeguards into Opus 5 intended to stop it from writing exploit code aimed at real targets. According to the researchers’ own account, they worked around those protections by disguising their own test server as a capture-the-flag practice environment, then let the model iterate on the exploit in an automated loop. They’re clear that this wasn’t a fully autonomous hack — a skilled human was directing the process throughout — but the jump in capability between the two model versions, on the exact same problem, is what has drawn the most attention from other security researchers.
Once inside a compromised OpenAI employee’s Codex account, the researchers found it was connected to OpenAI’s internal GitHub organization. Rather than reading any proprietary source code, they had Codex open a single, harmless pull request in OpenAI’s internal repository purely to prove the level of access was real, then stopped testing entirely.
How OpenAI Responded
Hacktron says it reported the forum vulnerability to OpenAI through its Bugcrowd bug bounty program on July 25, and that OpenAI confirmed a fix the same day. The company later paid the team a $6,500 bounty on September 1, specifying that the award recognized the OpenAI-side finding rather than the underlying Discourse vulnerability, which fell outside OpenAI’s bounty scope.
In a statement to VentureBeat, an OpenAI spokesperson confirmed the response, saying the company had “narrowed the permissions on Community sign-in tokens and revoked affected tokens and sessions.” OpenAI has not published its own detailed writeup of the incident.
What’s Still Unverified
Hacktron has framed the OpenAI case as one piece of a larger project it calls “HEIF Heist,” claiming the same class of image-decoding bugs turned up in software used by other major platforms, including Slack, Meta products, GitHub Enterprise, and the Next.js web framework. Reporting on this broader campaign has been more cautious: the Next.js flaw has been independently confirmed through Vercel’s own advisory, but the wider claims about other companies have not been documented with the same level of technical detail or independently verified by the affected companies. Readers should treat those broader claims as reported by Hacktron, not as independently confirmed facts.
There is also no public evidence that the OpenAI vulnerability was ever exploited maliciously outside of this authorized research, and as of mid-September it did not appear on the U.S. government’s catalog of vulnerabilities known to be under active attack.
Why It Matters
The most important takeaway isn’t really about OpenAI specifically — it’s about what happens when AI coding agents get connected to real business systems. As companies give AI tools access to source code, email, chat platforms, and cloud storage, an AI account effectively becomes an identity with real permissions attached to it. If that account is compromised, whatever it’s connected to becomes reachable too.
The capability jump between Claude Opus 4.8 and Claude Opus 5 on the exact same exploitation task is also a preview of a broader trend: frontier coding models are making types of technical security work — like turning a memory-corruption bug into a working exploit — dramatically faster than it used to be, for both defenders and attackers. Anthropic itself has previously reported that criminal and state-linked groups have already attempted to use Claude models for real intrusions, not just research. For a deeper look at how Claude compares to other leading AI assistants, see our ChatGPT vs. Gemini vs. Claude comparison, and for more on Anthropic’s recent momentum, see our coverage of Anthropic’s second profitable quarter.
Practical Takeaways
- Businesses running their own Discourse forums should confirm they’re on a patched release (2026.7.0, 2026.6.1, 2026.5.2, or 2026.1.6) and rebuild rather than assuming a web-interface update alone fixed it.
- Any organization sharing single sign-on between a public-facing service and internal tools should treat that shared trust boundary as a serious risk, not a convenience.
- Companies connecting AI coding agents to source repositories, email, or collaboration tools should treat those AI accounts with the same access controls and monitoring as privileged human accounts.
Conclusion
This incident was contained, disclosed responsibly, and fixed within hours — which is exactly how this kind of research is supposed to work. But it’s also a concrete, verified example of something security researchers have been warning about all year: AI coding agents are closing the gap between finding a vulnerability and actually exploiting it. As more companies connect AI tools to sensitive internal systems, the identity and access boundaries around those tools — not just the AI models themselves — are becoming the thing that actually needs defending.
Sources
- VentureBeat — “OpenAI hacked by small team of white hat security researchers using Anthropic’s Claude Opus 5” (September 17, 2026) — venturebeat.com
- The Hacker News — “Claude Opus 5 Helped Researchers Take Over OpenAI Staff Accounts via Chained Flaws” (September 19, 2026) — thehackernews.com
- Discourse (GitHub Security Advisory GHSA-vhm9-85gw-x335) — “RCE via malformed HEIF file” (July 28, 2026) — github.com/discourse

1 thought on “Claude Opus 5 Helped Security Researchers Breach OpenAI Staff Accounts”