
Researchers Used Anthropic’s Claude to Hack Into OpenAI: Here’s What Happened
In a striking demonstration of how AI is reshaping cybersecurity, a three-person security research team at startup Hacktron AI used Anthropic’s Claude to break into OpenAI’s internal systems, chaining two critical vulnerabilities to gain access to employee ChatGPT and Codex accounts. The incident, first reported by The Wall Street Journal and detailed in TechCrunch, unfolded over less than 72 hours in July 2026 and earned the researchers a $6,500 bug bounty from OpenAI.
The attack wasn’t malicious. Hacktron carried out the operation as part of OpenAI’s bug bounty program and reported its findings responsibly. But the implications are far-reaching: off-the-shelf AI tools can now be weaponized to find and exploit vulnerabilities in even the most sophisticated tech companies’ infrastructure.
As Matt Fredrikson, CEO of AI security firm Gray Swan, told TechCrunch: “For $200 a month, anyone can use these tools and hack into a company like OpenAI. If it can happen to them – and I don’t think they’ve been slouching recently on cybersecurity hygiene – it could happen to anyone”.
How the Hack Unfolded: From Image Upload to Internal Repos
The Entry Point: A Mundane Image Upload
The researchers found their way into OpenAI on July 25, 2026, via a flaw in Discourse, the third-party software powering OpenAI’s community forum at community.openai.com. The entry point was surprisingly ordinary: an image upload.
When users posted HEIF or HEIC image files – the format iPhones use by default – to OpenAI’s forum, Discourse passed them through a chain of behind-the-scenes tools to convert them into standard JPEGs. The first stop was ImageMagick, a decades-old open-source utility for resizing images. Because ImageMagick’s usual toolkit couldn’t handle Apple’s format, it handed the file off to another library called libheif to do the decoding.
Buried inside libheif was a heap buffer overflow vulnerability. A specially crafted image caused the library to miscalculate where one image was positioned on top of another – enough to hijack the server. This vulnerability, tracked as CVE-2026-32882 with a CVSS score of 8.8 (High), had actually been fixed upstream months earlier, but the fix was never formally flagged as a security vulnerability, meaning it never received a CVE number and was never backported to Debian.
Read also:
- YouTube’s New Custom Feeds Feature Lets You Build Your Own Algorithm With AI
- Google’s Gemini AI Hacked Three Companies During a Security Test
- Jensen Huang Says GPT-6 Astra Was Trained on 100K Nvidia GPUs, 400K More Coming
Claude Opus 4.8 Struggled – Then Opus 5 Changed Everything
What makes this story particularly compelling is the role of AI model capabilities. The researchers initially used a special version of Claude Opus 4.8, made available to qualified cybersecurity practitioners, to develop the exploit. But it couldn’t build a working exploit against Discourse’s default configuration with ASLR (Address Space Layout Randomization) enabled.
That changed overnight when Anthropic released Claude Opus 5. As Hacktron wrote in its detailed blog post: “Opus 4.8 struggled across several sessions to produce a working exploit. Within hours of Opus 5’s release, we gave it the same problem and it succeeded”.
According to Hacktron’s timeline, Opus 5 produced a working ARM64 exploit for a local Mac within three hours. By 6:00 a.m. on July 25, the team had confirmed local remote code execution (RCE) through an image upload. They then placed Claude in an autonomous goal loop against their own Discourse Cloud instance to refine the exploit before deploying it against OpenAI’s actual instance.

The Second Vulnerability: SSO Misconfiguration
Once inside the Discourse server, the researchers found another flaw – an SSO misconfiguration in OpenAI’s identity infrastructure. This allowed them to take over users’ ChatGPT and Codex accounts, including those belonging to OpenAI employees.
“We then took over an OpenAI employee’s account, whose Codex was connected to OpenAI’s GitHub organization,” Hacktron wrote. To demonstrate impact without actually accessing any internal code, they used the employee’s Codex to open a harmless proof-of-concept pull request in OpenAI’s internal monorepo.
The entire operation – from initial discovery to access to OpenAI’s private repository – took less than 72 hours.
OpenAI’s Response and the $6,500 Bounty
Hacktron reported the vulnerabilities to OpenAI and Discourse immediately. OpenAI confirmed the issue was fixed roughly 14 hours after the initial submission. The researchers also reported the libheif flaw to Discourse through HackerOne, and Discourse had a fix ready by Monday, adding image-processing sandboxing as defense in depth. Discourse published a security advisory with patch and rebuild guidance.
OpenAI paid the team a $6,500 bounty, though it clarified that testing against the Discourse-hosted community forum was explicitly excluded from its bug bounty program – the award recognized the OpenAI-side finding, not the actions against Discourse.
OpenAI says it has resolved the issues Hacktron uncovered.
Why This Matters: AI Is Collapsing the Expertise Barrier
The incident highlights a fundamental shift in the cybersecurity landscape. As Hacktron founder Mohan Pedhapati put it on X: “AI is reducing the amount of scarce expertise needed to develop exploits. Work that once took months can now take days”.
Hacktron’s broader research project, called HEIF Heist, traced the libheif vulnerability across Slack, Meta, GitHub Enterprise, Ruby on Rails, and Node.js frameworks such as Next.js, Astro, and Gatsby. The entire two-month campaign cost less than $3,000 in tokens and was conducted by just three researchers. Adapting the exploit to each new company typically took only one or two days.
“We are not aware of any company that detected the activity except Shopify, even after thousands of images were sent and their image processors repeatedly crashed,” Hacktron noted.
The Bigger Picture: AI Security Under Pressure
This incident comes at a moment when top AI companies are under growing pressure over safety. It follows several weeks after OpenAI’s own AI agents broke containment during a cybersecurity evaluation and hacked Hugging Face, demonstrating just how capable AI models are getting at making their own decisions.
The episode also puts a spotlight on where the line gets drawn for model capabilities. Claude Opus 5, the version that ultimately cracked the bug, hasn’t faced any security export restrictions – unlike newer version Mythos 5, which was temporarily locked down over concerns about its advanced hacking capabilities.
Meanwhile, open-weight models are increasingly catching up to the frontier in cyber capabilities. AI safety nonprofit SaferAI recently found that Chinese company Z.ai’s GLM-5.2 was only a few months behind OpenAI’s GPT-5.5 and Anthropic’s Claude Opus 4.7.
What This Means for Businesses
For organizations that process user-uploaded images – particularly HEIF, HEIC, or AVIF formats – the lesson is urgent. Hacktron’s warning is clear: “If your application processes user-controlled images and accepts .heic/.heif/.avif images, it is highly likely it is affected”.
The recommended defenses include:
- Update upstream: Install the latest security-patched libheif and libde265 packages. As of September 14, 2026, the latest upstream libheif security release is v1.23.4.
- Defense in depth: Disable untrusted HEIF/AVIF decoding where it isn’t needed, or isolate image-processing pipelines inside hardened, ephemeral sandboxes.
- Rebuild, don’t just update: For self-hosted Discourse, a web-interface update alone may not replace the underlying vulnerable Docker image. Run
git pullfollowed by./launcher rebuild app.
The Bottom Line
The Hacktron incident is a wake-up call. As one AI pundit noted on social media: “[Hacktron] used Opus 5 to pull off the hack… The question that will be asked is, if these three guys can pull this off, what can a nation state do”.
AI is removing the protection that complexity once provided. Software has long benefited from a kind of security through obscurity – the code and even the vulnerability could be public, but turning a bug into a reliable exploit still required rare expertise, significant time, and knowledge of the target environment. AI is turning more of that scarce expertise into compute, and security assumptions must catch up.
For companies relying on open-source image-processing libraries, the message is simple: patch now, sandbox aggressively, and assume that AI-powered attackers are already probing your defenses.
Receive News Updates and Tutorials Through our Social Media Channels, join:
- WhatsApp: BloginfoHeap WhatsApp
- Facebook: BloginfoHeap
- Twitter (X): @BloginfoHeap
- YouTube: @BloginfoHeap



