Startups

AI guardrails frustrate cybersecurity researchers using frontier models

Security researchers told TechCrunch that OpenAI and Anthropic safeguards can block legitimate vulnerability work and push teams to local models.

Theo Nakamura

By Theo Nakamura · Staff Writer

· 3 min read

AI guardrails frustrate cybersecurity researchers using frontier models
Photo: TechCrunch

AI guardrails for cybersecurity researchers are creating a practical problem: safeguards meant to stop malicious hackers can also block legitimate security testing, researchers told TechCrunch. For anyone tracking the AI business, the tension shows that model access rules can shape how useful frontier systems are in high-stakes enterprise work.

Guardrails are restrictions built into AI systems to limit certain outputs, such as instructions that could help someone break into software. In cybersecurity, the hard part is that the same technical steps can be used to defend a system or attack it.

OpenAI and Anthropic both offer vetted access programs for cybersecurity users. OpenAI calls its program Trusted Access for Cyber, while Anthropic runs a Cyber Verification Program for some Claude users, according to the companies’ materials cited by TechCrunch. These programs are meant to give approved researchers fewer cybersecurity restrictions than standard users face.

The debate grew sharper after the U.S. government placed export control restrictions in June on Anthropic’s Mythos and Fable models, TechCrunch reported. The move was tied at least partly to a report alleging that the models’ anti-abuse controls could be bypassed for cyberattack activity. TechCrunch reported that export controls on Fable 5 and Mythos 5 were later lifted, with Fable 5 returning to general access on July 1 and Mythos 5 reintroduced only to vetted U.S. organizations during a government review.

How do AI guardrails affect cybersecurity researchers?

Security researchers said the restrictions can interrupt routine defensive work, especially when they ask a model to analyze whether a bug can be exploited. A zero-day is a previously unknown software flaw; an exploit is code or a technique that takes advantage of that flaw.

Chris Anley, chief scientist at NCC Group, told TechCrunch that having a model test whether a bug is exploitable can help defenders decide whether a vulnerability is serious enough to fix. He said prompts that ask a model to repair code can also reveal where the dangerous weaknesses are, which makes the same workflow both defensive and offensive.

When models refuse those requests, Anley said, his team may turn to open-source AI systems without comparable restrictions. Those models can be downloaded and run locally, meaning the user does not have to send sensitive code or vulnerability details to a cloud provider.

Paolo Stagno, chief technology officer at CrowdFense, told TechCrunch that his company uses frontier models for reverse engineering, which means studying software to understand how it works. He said CrowdFense avoids using cloud-based AI models to find vulnerabilities or build exploits because that could expose sensitive vulnerability data or place it into future training processes.

Giuseppe Cali, a researcher who finds zero-days and develops exploits, told TechCrunch that guardrails do not interfere with his offensive work because he does not use AI for that part. He said he uses AI to understand code and build supporting tools, while keeping bug discovery and exploit development in human hands.

Other researchers described a more frustrating experience. One researcher at a smartphone-component maker, speaking anonymously to TechCrunch because he was not authorized to talk publicly, said Anthropic’s tools were barely useful for vulnerability research because the company was not in Anthropic’s Cyber Verification Program.

Chris Thompson, CEO of RemoteThreat and founder of Offensive AI Con, told TechCrunch that restrictions inside OpenAI and Anthropic’s vetted programs can still behave inconsistently. He said researchers can end up spending time trying to get usable responses instead of analyzing vulnerabilities.

Thompson said that dynamic may push responsible researchers toward Chinese open-source models such as GLM, which can be run locally without vetting or usage limits. He urged frontier AI companies to broaden responsible access while holding abusive users accountable, arguing that defenders need capable tools as attacks become faster and more automated.

This story draws on original reporting from TechCrunch.

More from Startups

All Startups