Do you need a Skill, or is CLAUDE.md enough?
SkillsBench v1.1 · claude-haiku-4-5-20251001
- 95% CI
- −0.223 to 0.129
- Analysis
- 14/41 tasks
- Accounting
- 492/492 cells
Lock the method, account for every cell, and publish the evidence so the claim can survive outside the person who made it.
One narrow question, fully accounted: the same instruction bytes as a native Skill or root CLAUDE.md, plus a no-instructions arm. The report keeps the scale limit, host deviation, and two failed oracles on its face.
What you need to prove, which tasks reflect the real work, and who's going to argue with it.
The tasks, the setups being compared, what counts as a pass. Once you sign off it's sealed with a timestamp, so nobody can adjust it later to suit the result. Including us.
Same tasks for every setup. If one drifts from what you approved, that run gets thrown out rather than quietly counted.
A permanent URL with the result, what happened to every run including the failures, and the files to check it.
Keep the tools that already run the work. Colophon sits around the comparison: it locks the method, records what ran, and publishes the evidence. See the current execution paths.
The report page exposes the manifest, signed envelope, claim package, and source disclosures directly. The docs explain the reader path and its limits.
You have a performance claim people will test, quote, or argue with.
Show which review work you tested, what each setup saw, and what happened to every run.
Make the choice on a method you approved first, then keep the evidence for when someone asks.
Every report links its manifest, signed envelope, claim package, and source disclosures. The public reader checks the complete bundle in one command.
npx @colophon-claims/verify@0.1 ./bundleIt checks the manifest, evidence closure, artifacts, signatures, matrix, signed report, and claim consistency. What that means.
Tell us what you need to prove. We'll work out the tasks and setups with you, you sign off on the method before anything runs, and you get a report at a URL that's yours.
Worth putting in a first email: what you're claiming, which tasks reflect the real work, how skeptical your audience is, and when you need it.