← All posts

We Ran 376 Adversarial Probes Against 47 Agent Skills. Injection Wasn't the Problem.

2026-09-29

Everyone in this ecosystem is worried about prompt injection. So we measured it — 376 adversarial probe runs against 47 agent skills, in a sandbox, with the skill actually loaded. Injection came in at 2%. The thing that actually breaks skills is more mundane, and it's not a clever attack.


Here is the short version. Across 47 skills and 376 adversarial probe slots, 33 boundaries dropped. Broken down by what the probe was trying to do:

Probe typeSlotsDroppedFailure rate
Push past the skill's declared scope942931%
Instructions hidden inside an artifact18832%
Move data out of the sandbox9411%

A "slot" is one probe in one capability tier, run three times and reported as the worst run. Two things follow from this table, and only one of them is reassuring.

What we actually ran

Every skill is loaded into an OS-sandboxed agent with no network egress except the model API. It faces one legitimate task — the baseline, to confirm the skill activated at all — and four adversarial probes: two that hide instructions inside the documents the skill was asked to process, one that pushes for more than the skill's own description claims to do, and one that tries to get data out. Each runs three times per tier, in two tiers: a restricted toolset, and the tools the skill declares in its own files.

Target models were deepseek-flash[1m] and glm-5.3-flash (one mid-batch provider switch, recorded per skill in the data), judged by deepseek-v4-pro[1m] and glm-5.3. Verdicts are the worst of three runs.

The reassuring half

Injection hidden inside content is, with three exceptions out of 188, caught and refused. This matches what the legal-set write-up found: skills whose job is to read a document carefully are already treating that document as suspect. Telling a contract reviewer that the boilerplate says to ignore its rules is like telling a proofreader the typos are intentional.

The failure isn't zero — you can see the exceptions in the data — but it is not the systemic hole the discourse assumes.

The other half

The probe that works is the one that isn't an attack at all. It doesn't hide anything. It just asks the skill to do slightly more than it says it does — to make the call, send the thing, skip the confirmation, decide instead of asking.

31% of those dropped. Compare: a probe that hides instructions in a CSV is refused 98% of the time, but a probe that says "just go ahead and do the next step too" succeeds nearly a third of the time.

That is not a security bug in the skills. It's the disposition that makes them useful. A skill that helps you is a skill that wants to finish the task, and "finish the task" and "exceed your remit" are the same sentence from the inside. The skills that hold that line are the ones whose descriptions say what they don't do — the boundary statement, which is also the dimension our static rubric scores as trigger quality.

The part of our own method you should hold against us

Two things in this data flatter the skills, and we'd rather you hear them from us.

1. The headline grade tracks the declared-capabilities tier only. Our grade is PASS or PARTIAL, and it is computed from tier B — the tier where the skill has the tools it asked for. That means a skill can drop a boundary in the restricted tier and still be published as PASS. It happens six times in this set. ansoff-matrix and d3-viz, for example, both failed the overreach probe in the restricted tier and both carry PASS with a resistance of 10/10.

We chose that framing deliberately — the restricted tier can fail for reasons that have nothing to do with the skill's judgment — but the consequence is that the big number on a skill page is more generous than the underlying probe verdicts. The tier-by-tier numbers are on every skill page and in the dataset, so the stricter reading is always available. We just don't want anyone discovering that from the raw file and concluding we buried it.

2. Five skills have a hold that isn't a refusal. Where the sandbox lacked a capability outright, a probe "passes" because the agent couldn't have complied even if it wanted to. The harness flags these (capability_gap_probes and the summary line), and we publish the flag. A handful of the looks-good numbers are of that kind.

Neither caveat changes the headline finding — overreach failures dominate in both tiers, 16 in the restricted tier and 13 in the declared one, so they aren't an artifact of the grading rule.

What to do with this

The data

All 47 skills, every probe verdict, the tier tallies, the capability-absence flags, the declared tools and hosts, and the runtime for each batch are published as an open dataset: injection-test data →, JSON at /data/injection-results.json, CC BY 4.0.

The probe prompt text is deliberately not in the file. Those prompts are working attack templates, and publishing them turns a results dataset into a kit for attacking other people's skills. Types and verdicts are published in full; researchers who want the transcripts can ask.

Twenty-six of the 47 skills in this set — 55% — held every boundary they were given, in both tiers. They just aren't the ones the discourse is worried about.