← All posts

We Ran the Same Evaluation Twice. Seven of Eleven Scores Moved.

2026-10-10

Every score on this site comes from a model reading a SKILL.md against a fixed rubric. We re-ran eleven of them two hours later, same inputs. Seven came back different. Here's the data, because a score you can't audit isn't worth publishing.


What happened

We evaluated 21 newly imported skills, then re-ran the evaluation for two of the authors' repositories a couple of hours later (a fix to the retry logic meant re-running, not re-scoring by hand). Eleven skills were scored twice, from the identical SKILL.md, by the same model, against the same rubric:

SkillFirst runSecond runΔ
understand-explain8.39.0+0.7
presentation-styling7.38.0+0.7
understand-figma8.38.7+0.4
understand-onboard8.08.3+0.3
understand-knowledge8.38.0−0.3
understand-domain8.38.0−0.3
presentation-structure8.37.3−1.0
understand, understand-diff, understand-chat, vibe-to-agentic-framework——0

Seven moved. Four didn't. The largest swing was a full point, in both directions.

What this does and doesn't mean

It doesn't mean the scores are meaningless. The rubric has real signal: it consistently separates a skill that verifies its own output from one that doesn't, and every score ships with the rationale that produced it. The model agrees with itself on the large structural questions far more than it doesn't.

It does mean you should stop reading the decimals. A 8.3 and an 8.7 are the same finding. Anyone — including us — who presents a 0.2-point gap as a ranking is reporting noise. Treat the useful resolution as roughly half a point, and treat "in the 9s" versus "in the 7s" as the real distinction.

And it means the honest way to publish an automated score is with its rationale attached. You can read why presentation-structure got a 7.3 on the second run and decide whether you agree. A bare number gives you nothing to argue with.

One caveat about this specific experiment

Between the two runs we also changed the generation budget (the first run sometimes hit a token limit before answering), so a small part of the difference may be the harness rather than the model's judgment. We're telling you that instead of presenting a clean-looking table: the variance is real — we've seen it in other re-runs — but this particular measurement isn't a controlled experiment.

The cleaner statement of what we learned: if you re-run an LLM evaluation, expect movement, and design your publication around that — publish the reasoning, use coarse bands, and never put a 0.1-point difference in a headline.

Why we're publishing this at all

Because the alternative is worse. Every directory in this space produces scores with some model and some prompt, and almost none of them publish the rationale or measure their own variance. A number that looks precise and can't be audited invites exactly the wrong kind of trust.

The rubric and its weights are public. Every score's written rationale is on its skill page. If you think a specific one is wrong, you have everything you need to say so.