
⚡ TL;DR — Key Takeaways
- What happened: A new front in AI cyberattacks opened this week when Google launched Gemini 3.5 Flash Cyber, an AI model built to find and patch vulnerabilities faster than attackers can exploit them — the same week a Russian hacker’s jailbroken-Claude pentest tool made headlines.
- The defense side: Gemini 3.5 Flash Cyber found 555 unique confirmed vulnerabilities in testing, beating both Gemini 3.5 Flash (474) and Claude Opus 4.6 (363).
- The offense side: A threat actor known as “Trim” built “AI Pentest Checker,” a commercial tool combining jailbroken AI models with scanning tools to automate reconnaissance and vulnerability discovery.
- Access matters: Google is restricting Gemini 3.5 Flash Cyber to governments and trusted partners specifically because of its dual-use potential.
- The bigger picture: Both stories confirm what the Five Eyes warned about weeks earlier — AI cyberattacks and AI-driven defense are now evolving on the same timeline, in real time.
Table of Contents
For years, cybersecurity followed a predictable rhythm: defenses improved, attackers adapted, and the back-and-forth continued at a pace humans could mostly keep up with. That rhythm broke this past week.
Within 48 hours of each other, Google unveiled an AI model built to out-hunt vulnerabilities before attackers find them, and security researchers exposed a criminal operation using a jailbroken AI model to automate the exact same kind of hunting — for offense instead of defense.
AI cyberattacks and AI defense are no longer separate storylines; they’re now racing on the same clock, and the pace of that race is exactly what makes AI cyberattacks such a difficult problem to get ahead of.
What makes this moment different from previous AI cyberattacks milestones is the timing, not just the technology. Individually, an AI model that finds software flaws faster than humans, or a criminal tool that automates reconnaissance, would each be notable stories on their own.
Happening in the same week, from organizations on opposite sides of the same fight, they read less like coincidence and more like confirmation that AI cyberattacks and AI-driven defense have entered the same accelerating feedback loop — where every advance on one side almost immediately becomes a benchmark the other side has to match.
Google’s Answer: An AI Built to Out-Patch AI Cyberattacks

Google DeepMind released Gemini 3.5 Flash Cyber, a security-tuned model built into CodeMender, the company’s AI-driven vulnerability discovery and patching agent first unveiled in October 2025.
According to Google DeepMind’s own announcement, the model trades raw scale for speed and cost-efficiency, allowing CodeMender to invoke it up to five times per report so multiple sub-agents can analyze far more code paths in parallel than a single, expensive model call would allow.
The performance numbers are striking. In testing on the V8 JavaScript engine, per GBHackers, the model identified 555 unique confirmed vulnerabilities compared to 474 for mainline Gemini 3.5 Flash and 363 for Claude Opus 4.6 — including 101 issues both competing models missed entirely.
Google has already deployed the model internally across Chrome, Android, Cloud, Ads, and YouTube, and its Cloud Vulnerability Research team reportedly found remote code execution flaws in public APIs within just two hours of testing.
Crucially, Google isn’t releasing this model broadly. Per The Hacker News, Gemini 3.5 Flash Cyber will be exclusively available to governments and trusted partners through a limited-access pilot, with guardrails specifically configured to enable only defensive functions. That restriction is itself a tacit admission: a model this good at finding vulnerabilities is just as good at helping someone exploit them, an imbalance Google addressed head-on by keeping tight control over who gets access.
The Other Side: AI Cyberattacks Built From a Jailbroken Model

While Google was preparing its defensive rollout, a very different story was unfolding on Russian-language cybercrime forums.
According to Cato Networks’ research, a threat actor known as “Trim” first appeared in March 2026 sharing techniques to bypass Claude’s safety controls — creating benign context before a harmful request, reframing instructions to focus only on code structure, and retrying softened versions of previously refused prompts.
By June, Trim had turned those jailbreak techniques into a working product: AI Pentest Checker, a commercial tool that combines a manipulated AI model with established reconnaissance tools like Nuclei, ffuf, katana, subfinder, and Gitleaks to automate target discovery, endpoint enumeration, and vulnerability checks.
None of those underlying tools are malicious on their own — security professionals use them constantly — but wrapping them in AI-driven automation and selling access to other criminals is precisely the kind of AI cyberattacks the Five Eyes warned about in their June statement covered in our earlier piece on agentic AI cyberattacks.
Why This Matters More Than Either Story Alone
Taken separately, these are two ordinary tech-industry stories: a product launch and a threat intelligence report. Taken together, they’re a real-time demonstration of the exact dynamic security researchers have been warning about for months.
Google explicitly built safety guardrails into its newer Gemini 3.6 Flash model specifically targeting “cyber offense misuses,” according to SiliconANGLE — a direct response to the same jailbreaking techniques Trim was actively exploiting against a competitor’s model just weeks earlier.
Google even noted it benchmarked against Claude Opus 4.6 rather than newer competing models specifically because those newer releases perform worse at finding vulnerabilities — a direct result of improved safety guardrails limiting their offensive capability.
That’s a telling detail: the same safety measures that make a model harder to weaponize for AI cyberattacks also make it less useful even for legitimate defensive vulnerability research, illustrating exactly how difficult this dual-use problem is to solve cleanly.
What This Means Going Forward
Neither side of this race is standing still. Google plans to expand Gemini 3.5 Flash Cyber’s capabilities to include red-teaming and full enterprise defense over time, while tools like AI Pentest Checker are already being sold commercially to other criminals on underground forums — meaning the barrier to launching sophisticated AI cyberattacks keeps dropping even as defensive tools improve.
That combination is what makes this moment genuinely different from prior cybersecurity escalations: AI cyberattacks are becoming cheaper and more accessible to less-skilled criminals at almost exactly the same rate that AI-powered defenses are becoming more capable, which means the net advantage for either side is unlikely to stay fixed for long.
For organizations, the practical takeaway echoes what we’ve written before: patch faster, assume automated reconnaissance is already happening against your systems, and treat AI-generated vulnerability reports — whether from your own defensive tooling or a criminal’s toolkit — as a normal part of the threat landscape now, not a future risk.
Waiting for AI cyberattacks to become a bigger, more visible problem before investing in AI-assisted defense is no longer a reasonable strategy, given how quickly tools on both sides are already operating at scale. The organizations that treat this shift as already underway, rather than something still on the horizon, will be the ones best positioned when the next wave of AI cyberattacks inevitably arrives faster than expected.
The Bottom Line
The AI-vs-AI arms race in cybersecurity isn’t a metaphor anymore — it’s happening in real time, with named tools, named actors, and measurable benchmarks on both sides. Gemini 3.5 Flash Cyber proves defenders can move faster than ever at finding and fixing flaws.
AI Pentest Checker proves attackers are moving at the same speed, using the same underlying technology, sometimes stolen from the very companies trying to stop them. Whoever adapts fastest to that reality — not whoever has the better model on paper — will define how AI cyberattacks play out over the next year.
What’s striking about this particular moment in AI cyberattacks history is how symmetrical it is. Both sides are using variations of the same class of AI model. Both sides are measuring success in the same currency — unique vulnerabilities found per unit of time.
And both sides are learning from the other’s public disclosures almost as fast as they happen, which means the gap between an offensive AI cyberattacks technique surfacing and a defensive countermeasure appearing in response is shrinking to a matter of weeks rather than months or years. That compressed timeline, more than any single tool or benchmark, is the real story here — and it’s one that’s unlikely to slow down anytime soon.
Related: Norton Genie AI Scam Detector: Does It Actually Stop Scams in 2026? – Read the major features of the Norton Genie AI Scam Detector and understand its benefits in your daily life.
The Ultimate Guide to Machine Learning Threat Detection in 2026 – Machine learning threat detection catches attacks in real time instead of after the damage is done — but the same technology is arming attackers just as fast as it’s arming defenders.
McDonald’s Data Breach 2026: 64 Million Job Applicants’ Data Exposed Through AI Chatbot – Read major data breach in McDonald’s McHire Software.
FBI Data Breach 2026: Inside the Chinese Hack of a Secret Wiretap System – FBI Wiretap Network Breached: How hackers infiltrated one of the FBI’s most sensitive surveillance systems—and what it means for U.S. national security.
Frequently Asked Questions (FAQ)
Q1. What is Gemini 3.5 Flash Cyber and how does it relate to AI cyberattacks?
A security-tuned AI model built into Google’s CodeMender agent to find and patch vulnerabilities before attackers can exploit them — Google’s direct response to AI cyberattacks outpacing human-speed defense.
Q2. Where does the “555 flaws” figure actually come from?
It’s from Google’s benchmarking across a fixed set of evaluations (555 vs. 474 for mainline Gemini, 363 for Claude Opus 4.6) — separate from the narrower V8 JavaScript Engine test, where the same model found 55 issues.
Q3. What is AI Pentest Checker, and is it a real threat?
A commercial offensive tool built by threat actor “Trim,” combining jailbroken Claude models with real scanning tools — a genuine example of AI cyberattacks being productized and sold to other criminals.
Q4. Why is Google restricting access to Gemini 3.5 Flash Cyber?
Because of its dual-use potential — a model this good at finding vulnerabilities for defense could just as easily accelerate AI cyberattacks if misused.
Q5. Does better AI defence mean AI cyberattacks will become less common?
Not necessarily — both sides are advancing on a similar timeline, so this is viewed as an ongoing arms race rather than a lasting advantage for either side.
DISCLAIMER
This article is published for general cybersecurity awareness and educational purposes only. The information contained herein is based on publicly available threat intelligence research and media reporting as of May 2026. This content does not constitute legal, financial, or professional cybersecurity advice. Readers should consult a qualified cybersecurity professional for guidance specific to their situation. All external links are provided for informational purposes; AI Security Watch is not responsible for the content of third-party websites. The mention of any product, service, or resource does not constitute an endorsement.Disclaimer
