Benchmark publishing for agent configurations

Publish benchmark claims people can check.

Lock the method, account for every cell, and publish the evidence so the claim can survive outside the person who made it.

Read one

One narrow question, fully accounted: the same instruction bytes as a native Skill or root CLAUDE.md, plus a no-instructions arm. The report keeps the scale limit, host deviation, and two failed oracles on its face.

How it works

01

Tell us what you're claiming

What you need to prove, which tasks reflect the real work, and who's going to argue with it.

02

You approve the method before anything runs

The tasks, the setups being compared, what counts as a pass. Once you sign off it's sealed with a timestamp, so nobody can adjust it later to suit the result. Including us.

03

Everything runs the same way

Same tasks for every setup. If one drifts from what you approved, that run gets thrown out rather than quietly counted.

04

You get a report people can check

A permanent URL with the result, what happened to every run including the failures, and the files to check it.

Keep the tools that already run the work. Colophon sits around the comparison: it locks the method, records what ran, and publishes the evidence. See the current execution paths.

The report page exposes the manifest, signed envelope, claim package, and source disclosures directly. The docs explain the reader path and its limits.

Who this is for

You're shipping a skill, harness, or loadout

You have a performance claim people will test, quote, or argue with.

You're making a review-agent claim

Show which review work you tested, what each setup saw, and what happened to every run.

You're choosing between setups

Make the choice on a method you approved first, then keep the evidence for when someone asks.

Check it yourself

Every report links its manifest, signed envelope, claim package, and source disclosures. The public reader checks the complete bundle in one command.

npx @colophon-claims/verify@0.1 ./bundle

It checks the manifest, evidence closure, artifacts, signatures, matrix, signed report, and claim consistency. What that means.

Bring a claim

Tell us what you need to prove. We'll work out the tasks and setups with you, you sign off on the method before anything runs, and you get a report at a URL that's yours.

Worth putting in a first email: what you're claiming, which tasks reflect the real work, how skeptical your audience is, and when you need it.