say-it-straight against no skill

Every skill on every example

Eight model-written drafts, four sizes per language. Bars are median API seconds over three cached runs; whiskers are min and max. Rose is say-it-straight, blue is no skill, gray is everyone else. A hollow bar means the skill returned a status block instead of a rewrite. Each table row links to the same output in the comparison viewer.

Context each skill loads

Tokens added to the prompt when a skill is invoked, counted with Claude's own tokenizer. This is the fixed part of the cost; the variable part is thinking.

Effort is the largest single variable

The same Korean mid draft, no skill, one run per effort level. Every series number on this page was measured at high.

Method and validity

Full report with per-run spreads, rule findings, and reproduction steps: docs/superloopy-skill-benchmark.md. Raw per-run data: 2026-09-11-skill-benchmark-runs.json.