← All posts

We Scored 591 Agent Skills. Here's the Shape of the Data.

2026-10-10

A single score tells you about one skill. The distribution tells you where the ecosystem actually is — and which of the six things we measure it is worst at.


Every skill on this site gets a score out of 10, on a rubric with six weighted dimensions, with the written rationale published on its page. As of today that's 591 skills. Here's what the whole set looks like.

The distribution

BandSkillsShare
9.0+15926.9%
8.5 – 8.918731.6%
8.0 – 8.413022.0%
7.0 – 7.99716.4%
below 7.0183.0%

Two things worth saying about this shape.

It is not a normal distribution, and that is not a coincidence. The directory is curated: skills are added because someone looked at them first. A random sample of everything on GitHub that calls itself a skill would look far worse at the bottom — you can see the shape of that sample in the 3% tail we did let through.

The middle is thick. Between 8.0 and 8.9 sits 54% of everything. Which means the score is genuinely useful for ranking the top of a category, and close to useless for choosing between two 8.4s. If two skills differ by 0.3, our own evaluation noise can account for it (we wrote about that separately).

The six dimensions, averaged across everything

Here's the more interesting table. Same 591 skills, broken down by what we actually measure:

DimensionAverage
Trigger quality — does it say when to fire, and when not to7.74
Structure — SKILL.md vs references vs scripts7.88
Engineering — valid frontmatter, runnable code, resolvable paths7.94
Workflow design — steps, not topics; verification loops8.11
Content quality — concise, imperative, executable8.30
Security — what it reads, contacts, and can break9.57

The weakest dimension is the one that decides whether the skill ever runs. A skill that never fires is worth zero regardless of how good its instructions are, and at 7.74 the ecosystem is collectively mediocre at describing when to fire. This is the single most common failure we see — a beautiful 400-line workflow behind a description that says "helps with documents."

The strongest dimension is partly an artifact, and you should discount it. Security averages 9.57 for a structural reason: confirmed credential theft, backdoor, or exfiltration caps a skill's total at 39/100 and gets it delisted rather than scored low. The worst security performers are absent from this average, not counted as 2s. Read 9.57 as "the surviving population is well-behaved," not as "skills are secure."

What this is not

These are our scores, from our rubric, with our judgment calls. The dimension weights (trigger 20, structure 15, workflow 25, content 10, engineering 10, security 20) are a design decision, not a discovered truth — a rubric that weighted workflow differently would produce a different ordering. The rationale for every individual score is published so you can argue with a number rather than with the whole system.

The raw per-skill data behind the adversarial half of these evaluations is open. The rubric is public. If you compute these averages yourself from the skill pages and get something different, tell us — that means one of us has a bug.