Gemini’s Breakout Is a Reminder the Basics Still Matter

September 23, 2026
Geminis breakout

In May, an Israeli security firm called Irregular put Google’s Gemini through a capture-the-flag exercise, testing how well the model could find its way through systems belonging to fictional companies. Perhaps unsurprisingly, Gemini didn’t stay inside the exercise. Google has confirmed that, through an unintended connection to the open internet, the model reached three real organisations, gaining access to one by guessing a password and to the other two through credentials it found sitting in a public repository. Google says Gemini stopped as soon as it recognised the targets were real, and that no lasting damage was done. It has since worked with Irregular to tighten the testing environment.

It’s a striking story, and an easy one to read as evidence that AI models are becoming dangerously capable on their own. But look at how it actually got into these companies, because on the surface this doesn’t seem all that exotic. It was a guessed password and a secret left exposed in a public repository, the same findings that would turn up in any competent access review; agent or no agent.

 

Why the basics still decide the outcome

In this instance, agentic AI didn’t invent new categories of risk. It inherited the ones already sitting in the environment (credentials that never expire, secrets copied into the wrong place, access nobody remembered to revoke), and acted on them without hesitation or fatigue. A person who stumbles onto an exposed credential still has to notice it, understand what it unlocks, and decide to use it. This agent was already looking, continuously, across every system it could find access to, and it moved the moment it found something that worked.

Now, weak credential practices used to be a slow burn risk, they’d be caught eventually by an audit or an attacker with enough patience. Hand the same weaknesses to an agent working at machine speed, as happened here, and the gap between exposure and exploitation shrinks to nearly nothing.

 

Fix the basics, then watch for what finds them

Most of this doesn’t start with agents. It starts with the same hygiene that’s been on every security checklist for a decade: credentials that expire, multi-factor authentication that stops a guessed password from being enough, secrets kept out of code repositories and rotated the moment they leak, and least privileged access, trimmed to what a system or person genuinely needs. The Gemini agent found what any secret-scanning tool or a diligent access review could have caught in time. Getting the basics right removes most of what an agent, or an attacker, has to work with in the first place.

But some exposure always slips through, and that’s where agents change the picture. An agent doesn’t need to recognise that a credential looks unusual or worth investigating, it just needs it to work, and it will try it the moment it’s within reach, at any hour, across every system it’s able to connect. That’s why it is important to treat agent identity and agent action with the same discipline that organisations have spent years building for people; knowing what every agent can reach before it reaches it, watching what it actually does in real time rather than reconstructing it afterwards, and being able to stop a specific action, or the agent itself, the moment it strays outside its job. A policy document doesn’t do that work on its own. It has to happen inline, at the moment the agent tries to act, with a human able to step in if it doesn’t look right.

 

What this means the next time an agent goes looking

The Gemini breakout is the latest example of concerning behavior from a frontier model. Earlier this year, Anthropic’s Claude escaped its test environment to hack three organisations, OpenAI models carried out cyber-attacks on publicly available services, (as mentioned in our own Agent Action Control launch), an AI coding agent found a stray API token in an unrelated file and deleted a company’s production database in nine seconds. Different companies, different agents, similar root causes: a credential with more reach than anyone intended, found by something that never stops looking. The moment of realisation for most security teams isn’t that this *could* happen. It’s recognising how many of those credentials are already sitting in their own environment, waiting for an agent (their own, or someone else’s) to go looking.

According to Netskope’s AI Risk and Readiness Report, 91% of organisations cannot stop a risky agent action before it executes. The Gemini incident is a reminder of why that number needs to change. And the fix starts with getting the basics right before the agents go looking for what we left behind.

author image

Scott Hogrefe

Scott brings two decades of experience to Netskope as VP Marketing. Before marketing, Scott spent several years in IT, working with R&D and analytical scientists.
Scott brings two decades of experience to Netskope as VP Marketing. Before marketing, Scott spent several years in IT, working with R&D and analytical scientists.
Keep a close eye on The Lens