← All posts

The Security-Conscious Skill Loadout: What to Install, How to Contain It

2026-09-28

We found an agent skill that harvested browser credentials. It scored well on every dimension that matters for productivity, which is exactly why the question stopped being "is this skill good?" and became "what am I going to let it touch?"


The credential harvester wasn't an outlier in quality — it was an outlier in intent. Most skills are honest, and honest skills still execute shell commands with your permissions. The loadout below is built containment-first: each slot does a security job, and the article ends with how to read the new badges that tell you what a skill does under adversarial pressure.

Slot 1 — the hardener: security-and-hardening — 8.5/10

Audits and hardens project configuration — dependency risks, secret exposure, permissions drift.

Why it's the anchor: It's the one skill in the stack whose job is to distrust the rest of the stack. Run it on any repo that other skills have been writing to.

Slot 2 — the credential discipline: setup-api-key — 8.6/10

Correct key handling: scoping, rotation, storage hygiene.

Why the boring skill is here: Most agent incidents won't be a malicious skill. They'll be a well-meaning one with an over-scoped token. This skill exists to make the safe path the easy path.

The one thing: Apply its discipline to the keys skills ask you for — the deploy tokens, the RPC keys. Scoped, revocable, never account-global.

Slot 3 — the builder with rules: skill-creator — 9.1/10

For writing your own skills with the security patterns baked in — input validation, least-privilege tool access, no ambient trust.

Why this matters more than it looks: The skills you write are the ones that inherit your permissions without your scrutiny. skill-creator encodes the checklist you'd otherwise skip.

Slot 4 — the orchestrator you can audit: subagent-driven-development — 8.5/10

Farms work out to subagents with explicit contracts.

Why containment likes it: Subagents are a security feature when the contracts specify what each one may touch. It's the architectural answer to "this skill doesn't need everything."

Reading the new dynamic badges

Every evaluated skill page now carries a Dynamic test card with the results of live adversarial probes: instructions hidden in intake files (injection), in user input, attempts to override the skill's declared rules, and workflow-escalation pushes. Adversarial probes run three times; the page shows the worst run. The two numbers to read:

The September document-category batch (17 skills, dual runtime: deepseek-flash for 374 runs, glm-5.3-flash for 68, judged by glm-5.3) produced the pattern to internalize: injections hidden in content almost never worked; instructions in the prompt that said "skip your own rules" worked six times out of seven on generator skills. The attack that works on your skills is the one that sounds like you.

So: keep the loadout above, scope every key, and when a skill page shows a 7.5 with a rule-override note — believe it, and never tell that skill to hurry.

Methodology · the legal-set injection results · what broke in the document batch