More than half of patches generated by artificial intelligence are flawed, introducing new vulnerabilities or failing to fix existing ones. Research on two prominent commercial models, OpenAI's ChatGPT 5.5 and Anthropic's Claude Opus 4.8, revealed that their automated patching capabilities are more likely to create exploitable patches than provide effective fixes1. This raises significant concerns about the expanding attack surface as AI-generated code becomes increasingly prevalent. The study's findings suggest that the use of large language models to generate patches may actually increase security risks, rather than mitigate them. As AI-generated code continues to proliferate, the potential for malicious hackers to exploit these vulnerabilities grows. This matters to security practitioners because it highlights the need for rigorous testing and validation of AI-generated patches to ensure they do not introduce new security risks.
More than half of AI-generated patches are broken
⚠️ Critical Alert
Why This Matters
LLM developments from OpenAI reshape both capability and risk surfaces — security implications trail the hype cycle.
References
- CyberScoop. (2026, August 7). More than half of AI-generated patches are broken. CyberScoop. https://cyberscoop.com/ai-code-patching-security-risks/
Original Source
CyberScoop
Read original →