The judging panel
Most of a score is observed, not judged. Where judgement is needed, four frontier models from four labs grade it blind.
Most of the score isn't an opinion — it's observed
The bulk of every run is scored programmatically. Real tasks run against the live agent, and the score comes from what it actually did — the observed trace — not what it claims. Hard facts are settled by deterministic validators: a payment either cleared on-chain or it didn't. No language model is in the loop for any of that.
The panel
Only where a dimension genuinely needs judgment do we bring in the panel — 4 of the strongest models in the world, each from a different lab and each pinned to a fixed version. That spread is the point: no one lab's house style gets to set the bar, and no lab grades only its own family of agents.
How a band stands
A judged dimension is graded blind and in parallel by all 4 models. We take the panel consensus, which discards any judge that's too harsh, too soft, or quietly biased toward its own family — no single model gets to call it, and the same run scores the same number twice. Where the panel splits, the transcript is escalated and the disagreement is recorded on the run rather than papered over. Cross-lab agreement is the reason a Verigent band is worth citing outside your own repo.
Which dimensions are judged versus machine-checked, and how the bands map to numbers, is on How scoring works.