The 20-Point Field Most Agent Skills Get Wrong
A skill that never fires is worth nothing, no matter how good its instructions are. And across 591 scored skills, the field that decides whether it fires is the one the ecosystem is worst at.
Every skill starts with a description in its YAML frontmatter. It is the only part of the skill loaded before the skill is chosen: your agent reads it, decides whether this skill matches what you asked for, and only then loads the rest.
That makes it a routing table, not documentation. It is also, by our scoring, the weakest of the six things we measure — 7.74/10 on average across 591 skills, against 8.30 for content and 8.11 for workflow design. Trigger quality happens to weigh 20 of our 100 points. It earns them being the hardest.
What the field has to do
A description that scores well does four jobs, and most skills do only the first:
- What it does — most skills do this.
- When to fire — common, usually phrased as
Use when the user asks to …. - When NOT to fire — rare, and the single strongest quality signal we see.
- The phrasings a person would actually type — rarer still.
- Marketing skills average 9.30 on this dimension.
- Finance & trading skills average 6.09.
- What does it produce? (A named artifact is better than a verb.)
- What would a user literally say to trigger it?
- What should route somewhere else instead?
- Is anything in here that is not about routing — a changelog, a philosophy statement, a roadmap?
The third one is where the money is. A skill with no negative space triggers on everything adjacent to it. Install three such skills and your agent starts loading the wrong instructions for ordinary requests, which is worse than not having them.
The same skill, two descriptions
Both of these describe an identical package that reviews a contract against a risk taxonomy:
Weak:
description: Reviews contracts and identifies legal risks.
Strong:
description: Reviews a contract against a risk taxonomy and returns findings
with severity and suggested redlines. Use when the user asks to "review this
contract", "check this NDA", or "flag risky clauses". Not for drafting new
agreements (use an NDA generator), not for jurisdiction-specific compliance
questions (use a compliance skill), and not for summarizing a contract without
risk analysis.
The second one is longer, and length is not the point — the exclusions are. It tells the router three things that should route elsewhere, which is what keeps the skill from firing on the 80% of adjacent requests it would handle badly.
What the data says about the gap
Trigger quality is not evenly distributed across the directory:
A 3.2-point spread is enormous for the same rubric item. But before you conclude that marketing people are better at naming things: that number is mostly a story about authorship. The marketing set is dominated by a small number of coherent, well-documented skill families; the finance set has grown by accumulation from many independent authors shipping crypto and market-data tools. The categories differ far less than the people writing in them.
A five-minute fix
Open your skill's frontmatter and check whether the description answers, in this order:
If you only fix one thing, fix #3. It is the difference between a skill that is available and a skill that is useful.
The full rubric — including how trigger quality is scored out of 20 and what the written rationale looks like — is public, and every skill page shows its own.
