At Evoke, we monitor what AI agents do inside customer environments and flag risky actions that could lead to security incidents. We wanted to know whether a fast decision model, like Jev, could help us detect a wider range of risky behavior.
Jev, TypeSafe's System One model, answers yes-or-no questions you define about a piece of text in milliseconds, with a confidence score for each answer. It doesn't reason step by step or write text, so it's fast and cheap enough to run on every agent action.
As a quick first test, we ran 104,368 real tool calls from coding and chat agents through Jev and asked three questions of every call at once: does this look like a destructive action, a control bypass, or data exfiltration?
Could Jev pick out the risky ones? The short answer: yes.

Agents don't repeat themselves
Deterministic signatures work when the same command is always written the same way. Agents don't work like that. They write each command from scratch, so the same dangerous action can look different every time.
Take a schema drop. Here's the same destructive action, written four ways:

Container escapes work the same way. One agent uses the privileged flag. Another quietly mounts the Docker socket and describes the command as running "without privileged." Every variant needs its own signature, and the agent will eventually write one nobody thought of.
What we need is a check that asks what an action does, not what it spells. Large language models can do that, but they are too slow and costly to run on every tool call. Jev, which TypeSafe describes as a "smart if-statement," answers typed questions in70 to 500 milliseconds with a calibrated score. That lets us ask of every tool call: is this threat-shaped? In other words, does it have the form of an attack, whatever the intent?
What we tested
We scored 104,368 real agent tool calls and report aggregates only.
The unit is one tool call, judged from its own text. We asked Jev to recognize the shape of a risky action, not to decide whether it was authorized or where the data was going. Those questions are passed to deeper, context aware layers of Evoke's detection system.
We limited this initial experiment to three risk types, each grounded in the MITRE ATT&CK framework:
- Destructive actions: irreversible, broad deletion or overwrite of real data or infrastructure. Clearing node_modules does not count. Deleting a production database does.
- Control bypass: disabling or getting around a security control, such as turning off endpoint protection or the firewall, disabling audit logging, escaping a container, or skipping a required review.
- Data exfiltration: moving data out of the environment through a channel an attacker could use.
We measured two things:
- Catch rate: Of the calls that were actually risky, how many did Jev flag?
- Review load: Of all calls, how many did Jev flag for deeper analysis by our context aware system? At the volume agents operate, even a small share is too many for people to review by hand, so that next step has to be automated.
In addition, we deliberately sampled calls Jev scored low, because that is where misses hide. In 2,000 low-scoring calls, we found none that were classified as risky. That doesn't prove Jev never misses, but it suggests risky calls aren't slipping through in the part of traffic Jev considers safe.
Results
Jev caught 96% of risky calls at a 5.3% combined review load. Deterministic signatures on the same set flagged only a couple dozen calls, and most of those were false alarms. They caught just 3 of the 199 risky ones.

The combined review load is lower than the three added together, because some calls are flagged for more than one risk type. That combined number is the one that matters for cost, since everything flagged goes to the same deeper analysis. Tune Jev looser or stricter and you trade catch rate against it.
Most interestingly, we tested Jev's ability to see past obfuscation attempts. We tried various encoding techniques and planted known bad commands inside real, harmless ones, then disguised them with base64, hex, and eval wrappers. Jev caught the disguised destructive and bypass commands at the same rate as the plain ones. Harmless commands wrapped the same way were mostly not flagged, so Jev reacts to what the command does, not to the fact that it was encoded.
Threat-shaped is not the same as malicious
This experiment answered a narrow question, and the answer is yes: a fast model can flag threat-shaped tool calls across all three risk types, cheaply enough to run on every call. Three takeaways stood out.
- Static signatures alone don't scale with agents. Agents write every command from scratch, so a rule written for one form of an attack misses the next. Signatures still catch the patterns we know but can't keep up with the ones agents invent on the fly.
- Decision classifiers are a strong addition to agent detection. Agents take far more actions than any human operator. A model that answers in milliseconds makes it practical to check every one of them, not just a sample.
- Flagging an action is not the same as detecting a threat. Agents do threat-shaped things all the time as part of ordinary work. Deciding which ones matter takes two questions a single tool call can't answer:

Answering those questions takes detection built for agents. That is where Evoke focuses. It's why our4M rogue action detection approach is so effective. Threat-shaped calls become candidates for our context-aware system, which judges them against what the operator was actually trying to do and surfaces the ones that don't add up.
If your team runs coding agents, they are making threat-shaped calls every day. Talk to us to find out which ones actually matter.

