Rogue Agents: What's the Real Threat To Your Organization?

Jason Rebholz


Rogue agent summer is muddying the waters about where companies should actually be concerned. Most of the reports are taken out of context and distort the real issues. The issues that should scare you are getting the least attention.

Let’s recap the major rogue agent news:

  1. OpenFace (OpenAI + Hugging Face): The model found a zero-day vulnerability, developed an exploit to escape its sandbox, and then hacked Hugging Face to “cheat” at solving a Capture The Flag (CTF) test.




    The model’s safety guardrails were intentionally turned off to measure the model’s raw power.


  2. Anthropic: A misconfigured evaluation environment gave its models access to the Internet. This resulted in Anthropic’s models hacking three different companies. One where the agent hacked a real company whose name matched the fictional company name given in the CTF parameters and stole data. Another instance occurred when the model executed a supply chain attack by publishing a malicious Python package that 15 real-world systems mistakenly downloaded and executed before it was pulled. The last instance occurred after the model scanned roughly 9,000 systems connected to the Internet and used basic techniques to compromise a web application.


    As with the OpenAI situation, the model’s safety guardrails were intentionally turned off to assess the model’s power.


  3. Meta: The same evaluation environment misconfiguration that impacted Anthropic also impacted Meta’s Muse Spark 1.1 model. During model tests, it hacked an unnamed company via a vulnerability in a third-party service. Meta is still investigating, but the model safety card notes that safety guardrails are removed during capability testing, similar to Anthropic and OpenAI testing.


  4. UK AI Security Institute: Across 122 evaluations of two of AISI’s cyber challenges, AISI found 19 instances where AI agents powered by Mythos 5 and GPT-5.6 Sol took unsanctioned actions on the Internet. This included targeting real people and organizations. Actions included: submitting malicious code into a public open-source project, creating fake online identities to socially engineer the human maintainer, contacting real people to trick them into executing malicious payloads, planting hidden prompt-injection instructions targeting other AI systems, and leaving public messages offering to collaborate with other AI agents. Yeah, just that…



    The agents were intentionally granted Internet access for realism, and the models’ safety classifiers were turned off. The agent was never given instructions on what it wasn’t allowed to do. Not that prompt-level instructions are a reliable control, as they’re not enforceable.


  5. Moonshot: In another test that Frontier Security ran, Moonshot’s open-weight Kimi K3 model was allowed network access to GitHub to find the benchmark test’s repo and read the answers directly from the downloaded data. No hacks involved. It just found the shortest path.



    This was the stock open-weight model. Unlike the previous examples, there were no safety guardrails to turn off because they largely don’t exist beyond the minimal safety mechanisms built into the model (which can also be stripped). Such is the way with open-weight models.


  6. OpenClaw: The most recent example: someone using OpenClaw powered by Claude Opus 4.6 to make a gym reservation. When the user learned he was 4th on the waitlist, he asked OpenClaw if it was possible to move him to the top of the list. OpenClaw dug in and found that the gym’s reservation system didn’t check authorization for canceling other reservations. It tested this by removing the person in the first slot, bumping the user up to the third slot.



    This is an example where Opus 4.6 had safety guardrails enabled and did it anyway. Of all the examples organizations should worry about, this is the one. Because it’s the closest example to what your employees are already using. It’s a situation where the user never asked the agent to do anything malicious. It was just being overly eager to help and ended up committing a crime. And most companies don’t have the visibility to know when this happens, let alone the ability to stop it.


Did you pick up the trend there? Five out of six of those examples had no to limited safety guardrails in place. The same guardrails that come standard in the frontier labs (but not the open-weight models). In Kimi's case (the open-weight model), there was no hacking. The OpenClaw one…that’s one to pay attention to.

Here’s how I look at this. Frontier labs’ models are getting really good. But open-weight models aren’t far behind. The AISI puts open-weight models 4 to 7 months behind frontier labs, a tighter gap from 6 to 10 months through 2025.

In tandem, the frontier labs’ safety constraints, which keep the model in check, are improving. But these safety constraints don’t apply to open-weight models that attackers can use.

This leaves two primary outcomes in the next six to twelve months:

  1. External Threats: Attackers using open-weight models to automate cyberattacks. The capabilities shown in the OpenAI, Anthropic, and Meta cases are largely available in open-weight models. The preview of these rogue agents showed what the models are capable of when their safety guardrails are disabled. Under normal circumstances, those frontier lab models with safety guardrails off aren’t available to attackers. That’s where open-weight models can deliver.


    Key takeaway: Current models, both frontier labs and open-weight models, are good enough to hack companies. It may be messy, but it still works. For every organization, this means cybercriminals can and will automate their attacks. Plan for and expect an increase in the velocity of everything you’ve been defending against for the last two decades.


  2. Internal Threats: Your agent goes rogue. The OpenClaw scenario isn’t far off from what we’ve seen with other rogue agents deleting data. It’s an “innocent” mistake but one that has consequences. In our own research on agents going rogue, we’re seeing some really interesting (and terrifying) behavior. More on that to come in future posts. These issues stay small right up until they don’t. The more enterprise adoption and usage of Claude, Codex, Cursor, etc., the more people connect agents to more tools and data. Don’t expect agent-related mishaps to be gradual with that rate of adoption.


    Key takeaway: Start monitoring your agents. It may be difficult (or impossible) to keep them in line 100% today, but how is this different than your internal employees? Things will go wrong. You will have incidents involving agents in the next six months. Don’t wait for the damage to happen before asking if you have the right visibility and detection capabilities in place. Then start working your way to enforcing actions and nudging agents to safer paths. Here's a blog post on how we think about this at Evoke Security


Rogue agent summer may be coming to a close, and it’s getting darker in the morning. That doesn’t mean you have to lose sight of what your agents are doing. If you’re an organization that is meaningfully deploying agents today, the inside risk is your sleeper threat. The one that seems far away but suddenly pops its head up and makes for a very bad day. You don’t let your endpoints operate without EDR visibility. Why would you let your agents operate with zero visibility?

If you have questions about securing agents, let’s chat.

Your trusted partner in securing your agentic workforce.

2026 | Evoke Security Inc.

Your trusted partner in securing your agentic workforce.

2026 | Evoke Security Inc.