Close Menu
  • Home
  • AI
  • Art & Style
  • Economy
  • Entertainment
  • International
  • Market
  • Opinion
  • Politics
  • Sports
  • Trump
  • US
  • World
What's Hot

Live updates: US attacks Iran for 13th straight night

July 23, 2026

Fed interest rate decision: Soaring oil prices raise probability of interest rate hike

July 23, 2026

CXMT raises concerns about capital outflow ahead of blockbuster IPO

July 23, 2026
Facebook X (Twitter) Instagram
Smart Breaking News on AI, Business, Politics & Global Trends | WhistleBuzz
Facebook X (Twitter) Instagram
  • Home
  • AI
  • Art & Style
  • Economy
  • Entertainment
  • International
  • Market
  • Opinion
  • Politics
  • Sports
  • Trump
  • US
  • World
Smart Breaking News on AI, Business, Politics & Global Trends | WhistleBuzz
Home » How AI guardrails are hindering the work of offensive cybersecurity researchers
AI

How AI guardrails are hindering the work of offensive cybersecurity researchers

Editor-In-ChiefBy Editor-In-ChiefJuly 23, 2026No Comments6 Mins Read
Share Facebook Twitter Pinterest LinkedIn Tumblr Telegram Email Copy Link
Follow Us
Google News Flipboard
Share
Facebook Twitter LinkedIn Pinterest Email


For months, the AI ​​giant has devised special, vetted programs and strict guardrails to limit the use of its models by malicious hackers. However, these limitations currently impede the work of offensive cybersecurity researchers as well as legitimate network defenders.

In June, the US government imposed export control restrictions on Anthropic’s highly touted AI models Mythos and Fable. The move was prompted, at least in part, by a report that claimed it was possible to bypass model guardrails designed to prevent users from using the model to construct and execute malicious cyberattacks.

Regardless of whether this incident was truly motivated by fear of jailbreak, the fact is that Anthropic has repeatedly promoted Mythos as some sort of apocalyptic cybermachine that can only be made available to carefully vetted users, and with strict guardrails in place. (Export restrictions on Fable 5 and Mythos 5 have since been lifted. Fable 5 returned to public access on July 1. Mythos 5 was only reintroduced to vetted U.S. organizations as part of a government review process.)

This kind of gatekeeping is not unique to Mythos. Anthropic and Other Models and OpenAI both offer programs that give cybersecurity researchers access to models with fewer cybersecurity restrictions if vetted and approved: OpenAI’s Trusted Access for Cyber ​​Program and Anthropic’s Cyber ​​Verification Program.

These guardrails have been widely criticized, especially by researchers whose job is to discover unknown vulnerabilities in systems and devise ways to exploit them before criminals can attack them.

Mark Dowd, a prominent security researcher, said in a recent appearance on a cybersecurity podcast that “I’m not very comfortable with these random big companies making arbitrary decisions about what’s security-wise and what’s not.”

For decades, Dowd has been discovering “zero days” — previously unknown software flaws and exploits — and selling them to Western governments rather than reporting them to software manufacturers to patch them. Governments pay a premium for vulnerabilities because vulnerabilities that serve intelligence operations remain open.

Dowd acknowledged that his work can create bias, but he’s not alone. Several people involved in offensive cybersecurity, who actively probe systems for weaknesses, explained to TechCrunch how they use AI tools and address guardrails.

Chris Unley, principal scientist at security consulting giant NCC Group, said attempting to exploit bugs in AI models is an important step to confirming that they are real vulnerabilities worth fixing. But he said the guardrails can hurt defenders if the model refuses to fully answer the questions.

“This is where the whole attack and defense and guardrails part becomes important, because the prompt, ‘Fix this code,’ is not only an essential mechanism for defense, but it is also a roadmap for discovering critical vulnerabilities in your code base,” Anley said. “So the same tool is both an offensive tool and a defensive tool, and you can’t really choose between the two.”

It’s “like a hammer,” he continued. “You can’t build a house without a hammer. A hammer is definitely a tool, but it’s also a weapon.”

When he and his colleagues encounter such obstacles, they often turn to open-source AI models that have no guardrails.

Paolo Stagno, chief technology officer at Cloudfence, a well-known company that develops, acquires and sells unknown vulnerabilities to government agencies, agreed with Dowd, saying that with vetted programs and guardrails, AI companies are “basically treating their customers like children who need to be babysat.”

Stagno said he and his colleagues do use the Frontier model, but only for reverse engineering. He said they avoid using AI to find vulnerabilities or build exploits. Inputting that work into a cloud-based model risks exposing sensitive vulnerability data or absorbing it into future training runs. He said that step uses an open source model that runs locally because it doesn’t rely on data sharing outside the model.

Giuseppe Cali, a security researcher who discovers zero-days and develops exploits, said the guardrails have not hindered his work. That’s because he doesn’t use AI for offensive work. Instead, we use it for initial reverse engineering, understanding the code we’re analyzing, and building supporting tools. To that end, he said, AI tools speed up the process and allow them to focus on finding vulnerabilities.

“I still want to do the actual bug discovery and weaponization myself, and even if all the guardrails were lifted tomorrow, that wouldn’t change,” Cali said. “I’m jealous of my bugs, but I love this game too much to have a model play it.”

A researcher at a smartphone parts maker, speaking on condition of anonymity because he was not authorized to speak to the press, said his employer is not part of Anthropic’s CVP program, so the guardrails are too strict and the tool is of little use in finding vulnerabilities.

“When the wind blows, we do security-related things, and the wind stops and we can’t use it,” the official said.

Chris Thompson, CEO of cybersecurity company RemoteThreat and founder of Offensive AI Con, an event focused on offensive security and AI, said that from his experience with frontier AI models, guardrails are inconsistent and can behave differently every day. This is true even within the looser boundaries of Anthropic and OpenAI’s vetted programs.

“I think the practical impact is that you spend a lot of time negotiating with models instead of working on your core security program,” Thompson says. “Rather than analyzing vulnerabilities and reasoning through exploitability, we’re trying to find out why we’re getting inconsistent results or why the model over-sanitizes the output.”

As a result, researchers rely on or are pushed by Chinese open source models like GLM, which are freely downloadable models that can be run locally without review or usage restrictions, Thompson said.

“Responsible researchers are being forced out of U.S. government systems and into foreign-owned systems,” he said. “I think putting up these guardrails will do more harm than good.”

Thompson called on the AI ​​Frontier Institute to open up its programs, provide responsible access, and hold those who abuse its tools accountable, rather than further restricting them. Otherwise, he argued, defenders will lose the AI ​​race.

“There’s a big storm coming. There’s going to be a big wave of attacks at a speed and scale that we’ve never seen before,” Thompson said. “But those same security consulting firms and legitimate researchers who are trying to make a difference are now being suppressed.”

If you buy through links in our articles, we may earn a small commission. This does not affect editorial independence.



Source link

Share. Facebook Twitter Pinterest LinkedIn Tumblr Email
Editor-In-Chief
  • Website

Related Posts

OpenAI makes ChatGPT Health available to all users in the US

July 23, 2026

As generated media becomes crowded, Runway launches AI model router

July 23, 2026

AegisAI, founded by former Google security executive, raises $36 million to stop AI-powered spear phishing

July 23, 2026
Add A Comment

Comments are closed.

News

President Trump imposes new double-digit tariffs on dozens of countries | Donald Trump News

By Editor-In-ChiefJuly 23, 2026

As the current 10% tax expires, President Trump will impose new tariffs on 60 countries…

Ukraine’s Zelensky meets US far-right activist Laura Loomer Donald Trump News

July 23, 2026

US court grants release of pro-Palestinian scholar as legal battle continues | Donald Trump News

July 23, 2026
Top Trending

How AI guardrails are hindering the work of offensive cybersecurity researchers

By Editor-In-ChiefJuly 23, 2026

For months, the AI ​​giant has devised special, vetted programs and strict…

OpenAI makes ChatGPT Health available to all users in the US

By Editor-In-ChiefJuly 23, 2026

OpenAI today announced that ChatGPT Health, a feature that helps users with…

As generated media becomes crowded, Runway launches AI model router

By Editor-In-ChiefJuly 23, 2026

Runway no longer wants to be just an AI modeling company. We…

Subscribe to News

Subscribe to our newsletter and never miss our latest news

Welcome to WhistleBuzz.com (“we,” “our,” or “us”). Your privacy is important to us. This Privacy Policy explains how we collect, use, disclose, and safeguard your information when you visit our website https://whistlebuzz.com/ (the “Site”). Please read this policy carefully to understand our views and practices regarding your personal data and how we will treat it.

Facebook X (Twitter) Instagram Pinterest YouTube

Subscribe to Updates

Subscribe to our newsletter and never miss our latest news

Facebook X (Twitter) Instagram Pinterest
  • Home
  • Advertise With Us
  • Contact US
  • DMCA Policy
  • Privacy Policy
  • Terms & Conditions
  • About US
© 2026 whistlebuzz. Designed by whistlebuzz.

Type above and press Enter to search. Press Esc to cancel.