← All posts

How We Dynamically Test Agent Skills (Not Just Read Them)

2026-10-08

How We Dynamically Test Agent Skills (Not Just Read Them)

Reading a skill's code tells you what it claims to do. Running it tells you what it does. Our directory started with the first — a six-dimension static rubric. This year we added the second: dynamic testing with adversarial probes.

This is the methodology, end to end.

Why static review isn't enough

Our static security scan reads every script and checks for credential access, undeclared network calls, prompt injection. It caught a credential harvester with 60,000 stars. It works.

But static analysis has a blind spot: skills that behave differently at runtime. A skill whose code looks benign but changes behavior based on environment, time, or input. For that, you have to run it.

The testing pipeline

Step 1: Containment

Every dynamic test runs in a sandbox — no host filesystem, no network except through a logging proxy, no real credentials. The proxy records every outbound request. This is non-negotiable: you're executing untrusted code by design.

Step 2: Probe design

We don't just "run the skill." We send probes — crafted inputs designed to surface specific behaviors:

Step 3: Execution and recording

Each probe runs against the skill in its sandbox. We record: what the skill did, what it accessed, what it sent, where it tried to send it. Sessions and transcripts are kept — that's the evidence trail.

Step 4: Scoring

Results feed the scorecard. A skill that passes all probes earns its static score. A skill that fails a probe gets flagged, and the failure mode determines the response: disclosure gaps get noted, actual injection attempts get vetoed.

What we found so far

The headline result from our first large run — 376 probes against 47 skills: injection mostly didn't work against well-built skills. The failures clustered in skills with weak input handling, which is exactly what the static rubric already penalized. That's a good sign: static and dynamic signals agree.

But dynamic testing caught things static review couldn't:

The full dataset is open: /data/injection-results.json.

What this doesn't cover (honesty section)

Dynamic testing is probabilistic, not exhaustive. We probe known failure patterns; novel attacks won't be in our probe set. Sandbox escapes are theoretically possible. And we test the skill as shipped — a skill that downloads code at runtime is only fully testable at the moment we test it.

The right framing: static review plus dynamic testing plus a public methodology you can audit. Not a guarantee — a much better filter.


Want the full methodology? It's public. Every skill page shows both its static scorecard and, where testing has run, its dynamic results.