← All posts

"Which Skill Category Is Best?" Is the Wrong Question

2026-10-10

People ask which category of agent skill is worth installing. We have 591 scored skills across nine categories, so we can answer it — and the answer is that the question is the wrong shape.


Here's every category, with the average score and the actual range of skills inside it:

CategorySkillsAverageRange
Video & Motion438.726.5 – 9.7
Documents208.716.7 – 9.8
Developer Tools1198.685.7 – 9.8
Marketing1198.587.5 – 9.4
Writing358.516.5 – 9.4
Design & Creative368.405.5 – 9.5
Productivity1138.346.3 – 9.6
Data & Analytics308.176.0 – 9.4
Finance & Trading768.135.6 – 9.3

The comparison that matters

Best category to worst: 8.72 to 8.13 — a spread of 0.59 points. That is smaller than our own evaluation noise on a re-run (we measured swings of up to a full point on identical inputs).

The spread inside Developer Tools: 5.7 to 9.8 — 4.1 points. Inside Finance: 3.7 points. Inside Design: 4.0.

So the within-category variation is roughly seven times the between-category variation. Knowing that a skill is a "developer tool" tells you almost nothing. Knowing it scored 9.8 versus 5.7 tells you everything.

Why the averages look the way they do

The category averages aren't measuring categories. They're measuring who writes in them, and how concentrated that authorship is.

Marketing's numbers are one team's style. It has 119 skills and averages 9.30 on trigger quality — the highest of any category by a wide margin. It is also dominated by a small number of coherent, consistently written skill families. The high average is a fact about those authors, not about marketing as a domain.

Finance is many authors, unevenly. 76 skills, the lowest trigger average of any category at 6.09, and a 3.7-point internal range — the signature of a category that grew by accumulation, as independent people shipped crypto and market-data tools at different levels of care.

Category size makes the comparison worse. Documents has 20 skills; Developer Tools and Marketing have 119 each. A small category is one author's good week away from topping the table.

What to do instead

Stop filtering by category and start filtering by score and by reader. The two questions that actually predict whether a skill is any good:

  1. Does it verify its own output? The single strongest correlate of a high score across every category we've measured. Skills that check their work before declaring success cluster at the top; skills that say "done" and hope do not.
  2. Has anyone read its scripts? Which is not a category property at all.
  3. Categories are useful for browsing — if you need a spreadsheet helper, start where spreadsheet helpers live. They are not useful for ranking, because the thing you are ranking against is a distribution seven times wider than the difference between the groups.

    The per-skill scores and their written rationale are on every skill page; the collections are at collections if you want to browse by domain. The adversarial half of these evaluations — which skills actually hold their boundaries when an agent runs them — is open data.