Good news first. We made it another week without AI destroying humanity. Let’s take the wins where we can. Bad news now. More AI doomer news popped up.
- Anthropic CEO Dario Amodei released an essay warning about the risks of recursive self-improvement, where AI builds itself, thus accelerating development and making it hard for humans to understand what it’s doing. He then calls out the OpenAI and Hugging Face hack to highlight the threat of an agent swarm “taking over the entire internet with a persistent botnet” in the next 6-12 months. His three-step solution involves third-party evaluators monitoring progress, common global safety standards, and global pacing on development. Oh, and another plea for the US government’s help to stop Chinese distillation attacks against US AI companies. Is all of this feasible when global AI investment is forecast to exceed $1 trillion in 2026, and every country is racing to gain the next global advantage? Safety be damned.
- Security researchers at Hacktron found a vulnerability in libheif, a widely used software library that encodes and decodes HEIF and AVIF image formats. They used Claude to help build a proof-of-concept exploit in three hours, bypassing safety guardrails by telling Claude that the task was part of a CTF. They used that exploit against OpenAI’s forum, which ran on Discourse’s enterprise community platform, to gain remote code execution and find a misconfiguration in their Single Sign-On (SSO). This let them take over OpenAI employee accounts that, unsurprisingly, were connected to Codex. The researchers used the access employees granted Codex to submit code to OpenAI’s internal code repository.
- Google was fashionably late in disclosing that one of its Gemini models was also hacking into companies. They reported that during model evaluations, their model hacked into three companies by finding publicly available credentials and guessing website passwords. This was linked back to the same testing company that caused a ruckus for Anthropic and Meta. Google reported that they didn’t announce it sooner because the hack caused no damage after the model realized it was attacking real companies. It gives off vibes of a parent defending their 10-year-old child who punched a toddler on the playground for no reason. But hey, he stopped punching on his own, and the toddler's fine now.
It's easy to jump to apocalyptic outcomes when updates like these flood your news feed and the media amplifies everything. My take? The probability math doesn’t math.
Let me share research from my past that helps ground me. When I was a CISO at a cyber insurance carrier, I spent a lot of time working with data science and actuaries to assess the likelihood and severity of catastrophic events. These were months-long projects and for good reason.
Insurance carriers make an educated and data-driven bet on whether the premium they collect is worth the risk they take on. Our analysis followed the surge in ransomware attacks, the scare of self-propagating destructive malware like WannaCry/NotPetya, Russia’s invasion of Ukraine (which used destructive cyberattacks), and heavy dependence on hyperscalers to host core services for many companies. Suffice it to say, there was concern, and there still is.
One particular analysis stood out to me. The risk of a hyperscaler like AWS going down. For an insurance carrier, depending on how long the outage lasts, that could mean paying out money for business interruption. While I won’t pretend to be an expert in this area, in its most elementary form, think of this as recovering money for lost revenue due to the outage.
To assess this risk, we took an adversarial mindset. What would have to be true for AWS to go down for an extended amount of time, from hours to days to weeks? This is where I got a crash course in a key concept:
Compounding probability. This is the likelihood of two or more things happening together. Each step has to succeed, one after the other. For an AWS outage to happen for an extended amount of time, a lot of things must go wrong. We plotted this based on the steps an advanced threat actor would have to take to succeed. Each step became progressively harder for an adversary to accomplish, let alone avoid detection. By the end, we found the likelihood was only slightly higher than an extinction-level event from an asteroid hitting Earth.
We followed a similar methodology for other scenarios. The one with the worst outcome? A self-propagating destructive attack. I’ll spare the details, but the room went quiet when we saw how many years it would take to recover the losses. It far exceeded the lifespan of anyone in the room.
That quiet pause is where the world is today. But people are jumping to extreme conclusions without data, knowledge, or experience. Yes. AI is powerful. Yes. It will keep getting better, and faster. At some point, it will surpass human intelligence. Especially when we remember that humans make a lot of mistakes, too. I mean, honestly, have you met another human? We can be really stupid. But I digress.
The fact is, the current state is good enough to hack into companies. It’s also dumb enough to go rogue and hack companies when it shouldn’t, and it will hide the bad things it does from users. It will also follow a human’s orders to bypass internal security controls.
Does that mean AI will hack the world and take over the Internet in the next 6-12 months? Compounding probability says it's possible, but not likely. We aren’t defenseless on the cyber front. We just need to evolve our perspective on how to properly defend. Here are two ways I look at this:
Secure from the Outside In: Protect against scenarios where adversaries or rogue agents hack you. All attack playbooks are fair game, from new software zero-days to social engineering. Agents are already automating this, and you should expect and plan for increased velocity.
The good news is that the layered defenses you’ve been applying for years still work. An AI-enabled attack, just like a human-led one, has to have a series of things go right in succession to succeed. This still gives defenders an advantage. Yes, they will need some added muscle and some new layers, no doubt, but from what we’ve seen in how agents hack, it’s a bull in a china shop. Speed to detect, contain, and remediate is critical here. Every layered defense is breathing room for detection and containment. It’s how you stack the compounding probability in your favor.
Secure from the Inside Out: Protect the agents operating in your network. These are the enterprise agents you give your employees, like Claude, Codex, or Copilot. As we saw with Hacktron, they accessed a valid user’s account and used the access the victim gave their agent against them. I wrote about this in January 2026. As I called out previously, we’re seeing firsthand how employees use agents to bypass security controls, and we’re seeing agents as a new tool for insider threats. Don’t ignore the risks agents operating in your environment pose. They may not hack other companies, but they will cause your own security incidents.
Slowing down AI development isn’t a feasible response. But we can’t wait for a solution and do nothing on the cyber defense side. We have a brand new technology that is accessible to anyone with an Internet connection and a will to learn. Adversaries will absolutely use agents to proliferate their hacking campaigns. Defenders must use agents to counter the threat as it scales.
Yesterday the risks were paper cuts. Today, we’re dealing with pistols. But we’re not yet at a planetary-extinction-level risk.
If you need help securing your enterprise agents operating in your environment, that’s where Evoke can help. Let’s chat.

