Gemini" text with sparkling star icons on a geometric green and black background.

Google Unveils Gemini 3.5 Flash Cyber to Speed Up AI Vulnerability Hunting

Google launches an AI security model to find vulnerabilities faster and cheaper, challenging Anthropic and other rivals in cybersecurity.

In short

Google has launched Gemini 3.5 Flash Cyber, an AI security model built to find and patch vulnerabilities at lower cost than larger rivals. The company is first offering it to governments and trusted partners through CodeMender.

  • Google introduced Gemini 3.5 Flash Cyber as a cheaper AI security model for vulnerability discovery.
  • The model is being deployed first through CodeMender for governments and trusted partners.
  • Google says repeated calls to the model helped it find more issues than some larger rivals in testing.
  • The launch intensifies competition with Anthropic and Microsoft in AI cybersecurity.

Google has introduced Gemini 3.5 Flash Cyber, a lower-cost AI model designed to help security teams find and fix software vulnerabilities faster, first through its CodeMender coding agent for governments and trusted partners. The launch matters because it gives organizations a cheaper way to run repeated security scans at high speed, a capability Google says can uncover more flaws than larger, more expensive models.

The company says the new model is built on Gemini 3.5 Flash and tuned for cybersecurity work, including scanning code paths, identifying weaknesses and helping patch them. In Google’s telling, the goal is not to replace human defenders, but to let security systems explore more of a codebase than traditional tools can manage on their own.

The announcement arrives in a fast-moving race among major AI developers to dominate a niche but increasingly important market: AI systems that do not just write code, but also inspect it for flaws, exploit chains and defensive fixes. Anthropic, Microsoft and others have already been pushing similar tools, turning AI security into one of the most contested corners of the enterprise AI market.

What Google launched and why it matters

Google’s new offering is an AI security model aimed at one specific job: finding vulnerabilities in software quickly and at a lower cost than large frontier systems. Rather than positioning the model as a general-purpose chatbot or coding assistant, Google is packaging it as a specialized tool for defensive security workflows.

That positioning matters for two reasons. First, cybersecurity workloads can be expensive if they rely on large, general models that consume substantial compute. Second, defensive scanning often benefits from running many passes over the same source code, which makes cost efficiency as important as raw model strength.

Google says Gemini 3.5 Flash Cyber is intended to be a “cost-efficient and highly capable alternative” to larger AI security systems. In practice, that means it is meant to compete not only on accuracy, but on the economics of security automation.

How the model fits into CodeMender

Google says the model will initially be available to governments and selected trusted partners through CodeMender, the company’s security-focused coding agent. CodeMender can call the model repeatedly and at high speed, which allows it to inspect more code paths and search for additional vulnerabilities.

In other words, the model is not being released as a standalone consumer product. It is being inserted into an existing security workflow where its value depends on repeated, low-latency use.

Google says CodeMender can invoke Gemini 3.5 Flash Cyber many times quickly and at relatively low cost, giving the agent more opportunities to explore code paths and uncover security issues.

That repeated invocation is central to the product strategy. Security analysis often improves when a system can take multiple shots at the same codebase from different angles, and Google argues this model makes that affordable enough to do at scale.

How does Gemini 3.5 Flash Cyber compare with larger models?

Google says the new model can hold its own against much larger systems in benchmark testing, even when it is run several times. On the CyberGym AI cybersecurity benchmark, the company says Gemini 3.5 Flash Cyber delivered competitive results compared with significantly larger models when used up to five times.

That benchmark framing is important because Google is not claiming the model always outperforms the biggest models on a single pass. Instead, it is arguing that a smaller, cheaper model used iteratively can match or surpass more resource-intensive rivals in practical security tasks.

The implication is straightforward: in cybersecurity, the best tool may not be the largest one. It may be the one that can be run often enough, across enough potential attack surfaces, to surface the most useful findings.

Model Primary Role Cost Profile Notable Result Mentioned by Google
Gemini 3.5 Flash Cyber Defensive vulnerability discovery and patching Lower-cost, high-frequency use 55 unique confirmed issues in V8
Gemini 3.5 Flash Baseline model for comparison Lower-cost general model 47 unique confirmed issues in V8
Anthropic Opus 4.6 Higher-end AI security competitor More expensive than Flash Cyber’s approach 36 unique confirmed issues in V8
Anthropic Mythos 5 Compute-heavy security model Expensive, benchmark leader competitor Referenced as a larger model benchmark comparison

What did Google claim the model found?

Google points to two sets of results to show the model’s value. First, it says Gemini 3.5 Flash Cyber produced strong benchmark performance on CyberGym when invoked multiple times. Second, it says the model uncovered more vulnerabilities in the V8 JavaScript engine than the comparison systems cited in its post.

According to Google, the model identified 55 unique confirmed issues in V8, compared with 47 found by Gemini 3.5 Flash and 36 found by Anthropic’s Opus 4.6. Google also said the model found 10 issues that no other model detected.

That is a meaningful signal for a security tool, because bug-finding is often measured not just by total output but by the uniqueness of the findings. If a model can discover vulnerabilities that others miss, its value increases significantly for defensive teams looking for high-priority, previously unknown weaknesses.

Why the V8 results matter

V8 is Google’s open-source JavaScript and WebAssembly engine, used in Chrome and other environments. Because it is widely deployed and deeply complex, it provides a useful proving ground for vulnerability discovery tools. A model that can spot issues there can demonstrate relevance beyond a narrow synthetic benchmark.

The company’s emphasis on “unique confirmed issues” suggests it wants to show that Gemini 3.5 Flash Cyber is not simply rediscovering the same bugs as previous systems. Instead, Google is highlighting the model’s ability to push into unexplored parts of the codebase after multiple runs.

Why is Google targeting a cheaper security model now?

Google is moving now because the AI security market is getting crowded, expensive and strategically important. Security teams want tools that can help them hunt bugs at scale, but large models can be too costly for frequent, repeated analysis. A smaller specialized model can be easier to deploy and cheaper to run across broad codebases.

The timing also reflects growing pressure from competitors. Anthropic’s Mythos line has already set a high bar in the security-model niche, and Microsoft has reportedly used Anthropic’s technology for parts of its own vulnerability checks. Meanwhile, Chinese competitors are also claiming strong performance in the same space.

For Google, the opportunity is to show that it can compete not just with raw model size, but with efficiency and practicality. If the model can deliver nearly comparable security results at a lower cost, it could appeal to government and enterprise users who need to run repeated checks without blowing through budgets.

How Google’s move fits into the AI security arms race

The launch comes as large AI firms increasingly see cybersecurity as a flagship enterprise use case. Security is attractive because the value proposition is easy to explain: AI can help defend software faster than teams working manually, especially when codebases are large and vulnerabilities are subtle.

But the market is also becoming a proving ground for model quality. Vendors are no longer just demonstrating that their models can answer questions or write code. They are showing whether those models can autonomously reason through software, identify weaknesses and generate fixes that hold up under review.

That makes the field especially competitive. Anthropic has promoted Mythos 5 as part of its Project Glasswing initiative, describing it as a powerful but compute-intensive security model. Microsoft, for its part, has already leaned into AI-assisted vulnerability discovery, underscoring how rapidly these tools are being operationalized.

Anthropic’s Mythos 5 is presented as a high-powered security model within Project Glasswing, but Google is arguing that smaller, cheaper systems can be more practical for repeated defensive work.

What is Anthropic’s role in this story?

Anthropic is the key benchmark competitor because its Mythos family has become one of the best-known references in AI security tooling. Google’s launch is, in part, a direct response to that lead. By comparing Flash Cyber to Mythos and Opus models, Google is signaling that it wants a share of the same enterprise security market.

The comparison also highlights a strategic contrast. Anthropic’s approach has emphasized heavy-duty models that can be costly to operate, while Google is leaning into a more efficient model that can be used repeatedly in a broader workflow.

What the benchmark numbers suggest

The benchmark claims tell a story about iteration. One of the clearest themes in Google’s announcement is that performance improves when the model is called multiple times. That matters because security analysis is often a search problem rather than a single-response problem.

  • Repeated calls can expose different reasoning paths.
  • Lower cost makes large-scale scanning more feasible.
  • Multiple passes can uncover bugs a single pass misses.
  • Specialized tuning can outperform a general-purpose model in a narrow task.

This is also why the company is stressing that CodeMender can make many rapid requests to the model. The idea is to transform the model from a one-shot assistant into a continuous security probe.

If those results hold up in real-world use, the implication is that defenders could scan larger sections of code more often, potentially shrinking the window between the introduction of a vulnerability and its discovery.

Who will get access first?

Google says governments and trusted partners will be the first users of Gemini 3.5 Flash Cyber through CodeMender. That limited rollout suggests the company is taking a controlled approach, likely to reduce risk while gathering feedback from security organizations that can test the model in practical settings.

Restricting access initially may also reflect the sensitive nature of cybersecurity tools. A model that can spot vulnerabilities could, in the wrong hands, also help attackers better understand weak points in software. Companies introducing such systems often move carefully for that reason.

By starting with a smaller set of vetted users, Google can refine the system, validate its claims and assess how it behaves in real-world environments before a broader release.

Timeline of the AI security race

The following timeline shows how the market for AI vulnerability-hunting tools has intensified around major model releases and enterprise adoption:

Period Development Significance
Anthropic Project Glasswing Launch of the Mythos security model family Established a benchmark for powerful AI security tooling
Microsoft adoption Microsoft begins using Mythos for security checks Shows major enterprise appetite for AI-assisted vulnerability discovery
Google launch Gemini 3.5 Flash Cyber is announced Google counters with a lower-cost alternative
Initial rollout Release through CodeMender to governments and trusted partners Controlled deployment focused on defensive use

What this means for enterprise security teams

For enterprise security teams, the launch is a reminder that AI security products are moving from experimental demos toward operational tools. The appeal is obvious: the ability to scan more code, more frequently, at a lower cost.

Still, adoption will depend on trust. Security teams will want to know not only how many vulnerabilities the model can find, but how accurate those findings are, how often it produces false positives and how easy it is to integrate into existing developer workflows.

Enterprises also tend to evaluate these tools through a practical lens. A model that performs impressively on benchmarks may still need to prove that it can handle proprietary codebases, legacy systems and messy real-world dependencies.

What security buyers will likely watch next

Buyers are likely to focus on four questions before adopting any AI security model at scale:

  1. Does the tool find vulnerabilities that matter in production systems?
  2. Can it do so cheaply enough to run continuously?
  3. How does it compare with human security engineers?
  4. Can it be deployed safely without creating new risks?

Google’s launch is designed to answer the second question especially well. The company is betting that security teams will value a model that can be run often, not just one that looks powerful in a single demonstration.

Why the economics of AI security are changing

AI security has entered a phase where cost is becoming as important as capability. In many enterprise contexts, the best defense is not the most expensive model, but the one that can be used repeatedly enough to catch subtle bugs before they reach production.

That shift could reshape vendor competition. If customers can get useful results from a smaller model that costs less to operate, they may prefer that option over a heavyweight system that delivers only incremental gains.

Google is clearly making that argument with Gemini 3.5 Flash Cyber. The model’s value proposition is based on a simple equation: if a smaller system can search more broadly because it is cheaper to run, it may actually find more problems overall.

Bottom line

Gemini 3.5 Flash Cyber gives Google a new foothold in the emerging market for AI-powered security analysis, where lower cost and repeated use may matter more than sheer model size. With a first rollout limited to governments and trusted partners through CodeMender, Google is signaling both ambition and caution as it challenges Anthropic and other rivals in a rapidly escalating cybersecurity race.

If the company’s claims bear out in real-world deployments, the model could become an important tool for organizations that need to scan complex software constantly without paying frontier-model prices. For now, the launch makes one thing clear: the competition over AI security is no longer about who has the biggest model, but who can find the most bugs at the best price.

Frequently asked questions

What is Gemini 3.5 Flash Cyber?

Gemini 3.5 Flash Cyber is Google’s new AI security model built to help find software vulnerabilities and support patching work. Google says it is tuned for repeated, low-cost security scans rather than general chat or broad coding tasks.

Who can use Google’s new AI security model first?

Google says the first users will be governments and trusted partners. The model will be available through CodeMender, Google’s security-focused coding agent, before any wider rollout is considered.

How does Gemini 3.5 Flash Cyber compare with Anthropic’s models?

Google says Gemini 3.5 Flash Cyber is a lower-cost alternative that can deliver competitive security results against larger models, including Anthropic’s Mythos and Opus systems, especially when invoked multiple times during analysis.

Why is Google focusing on a cheaper security model?

Google is focusing on a cheaper security model because vulnerability hunting often requires repeated scans across large codebases. Lower operating costs make it more practical to run the model often, which can improve the chances of finding hidden flaws.

Share this 🚀