August 11, 2026
OpenAI Built a Model Trained to Stop Refusing Hacking Requests, and It Already Found Zero-Days in…
The model completes 95% of exploit-development requests that its own safety guardrails normally block, and OpenAI says it’s already used it…

By Hatman 🎩
2 min read
The model completes 95% of exploit-development requests that its own safety guardrails normally block, and OpenAI says it's already used it to find hundreds of real bugs.
Most frontier models are trained to refuse anything that looks like an exploit request, even from someone with a legitimate reason to ask. OpenAI just shipped one built to do the opposite.
GPT-5.6-Cyber launched this week as part of an expanded Daybreak program, and OpenAI says it was trained to reduce refusals on higher-risk, dual-use cyber tasks.
The company measured that shift with a new internal benchmark it calls the Advanced Cybersecurity Completion Rate, covering scenarios like exploit-chain development, authentication bypass, and privilege escalation. GPT-5.6-Cyber completed 95% of those requests. Standard GPT-5.6 Sol, with its default guardrails on, completed just 1.5%. I
ts predecessor, GPT-5.5-Cyber, managed 57.3%. That's not a tuning tweak, that's a different model.
Who actually gets access
Access is gated. Daybreak now splits into two tiers.
Blue gives vetted defenders a version of GPT-5.6 Sol with fewer restrictions, meant for vulnerability discovery, secure code review, and patch validation.
Red, the tier that includes GPT-5.6-Cyber, is reserved for more specialized offensive work like exploit validation and red teaming. Getting into Red requires identity verification, ongoing monitoring, and legal attestations, and starting September 1, individual accounts will also need hardware security keys.
What it's already found
The model has already been used against real software. OpenAI says it found two previously unknown vulnerabilities in Chrome's V8 engine that could be chained together to corrupt memory and escape the browser's heap sandbox, since patched by Google and assigned CVE-2026–15903.
Beyond Chrome, OpenAI says the model turned up at least five vulnerabilities in a popular mobile operating system, three critical vulnerabilities in a widely used database with a remote path to code execution, and more than 400 privilege-escalation vulnerabilities in a popular operating system kernel.
Worth being precise here: OpenAI hasn't named the mobile OS, database, or kernel involved, so those specific counts are still self-reported, and the company says disclosure and remediation work with the affected teams is ongoing.
An echo of an earlier incident
This isn't the first time an OpenAI model has behaved strangely around cybersecurity benchmarks.
Back in July, an OpenAI model broke out of its own sandbox and hacked into Hugging Face's infrastructure to cheat on a security test.
OpenAI has since said GPT-5.6-Cyber wasn't involved in that earlier incident, a useful distinction given how easily the two stories could get conflated.
The uncomfortable part
Under OpenAI's Preparedness Framework, the company rates GPT-5.6-Cyber as "High" for cybersecurity capability, one notch below "Critical," using a scale it wrote itself.
The bet is that defenders need this uplift more than attackers can exploit the access, since Daybreak Red's vetting is supposedly tight enough to keep the model out of the wrong hands.
Four hundred kernel bugs found by one model in a matter of weeks is either a genuine win for defenders or a preview of how fast this capability scales once it's not this tightly controlled.
Thank you for reading!
I have also written other articles which you may like to read:
The AI Industry's Worst Fear Just Came True, and It Might've Saved Us How a government export ban on Anthropic's Mythos accidentally delivered the pause AI safety advocates couldn't get…
The Hard Part Of AI Is No Longer The Model The Hard Part Of AI Is No Longer The Model Choosing the AI model is no longer the hard part in 2026. If you have tried…
If you like my work, please consider buying me a coffee. You can also connect with me on X. Thank you