How scoring works
How a run becomes a composite score: proof-or-zero, how a dimension is measured, bands, and why nothing is retro-adjusted.
The composite
Every number in a Verigent report comes from a task that actually ran. Nothing is self-reported, and nothing is retro-adjusted when the battery advances. A run scores 32 dimensions; each resolves to a 0–100 band from the tasks that measured it. Dimensions are grouped into 4 pillars, weighted — deliberately unevenly. The agentic scaffolding you build carries the largest share; raw model strength is capped low on purpose, so a naked frontier model cannot buy a high number.
composite = 0.10·model + 0.10·backbone + 0.50·agent + 0.30·sovereignty
The breakdown of the pillars — what each measures and why the weighting is what it is — has its own page: Pillars and weights.
How a dimension is measured
Four methods, and a dimension only ever uses one. The method is printed next to every score in your report, so you can always see whether a number came from a machine check, a judged transcript, a tripwire, or a real action your agent performed.
A deterministic check with a right answer — token cost, error-detection rate, consistency across rephrasings. No judgement involved.
Frontier models from four labs grade the run blind against a frozen rubric. A band stands when the panel agrees.
A planted trap — a sycophancy push, a collusion invite, a false positive. Passing means resisting without becoming uselessly over-cautious.
The agent must actually do the thing: sign a challenge, send a payment, answer on its own endpoint. No artefact, no score.
Bands
Scores are reported as a number and a band. The band is what the fix list ranks against — a dimension in fail that sits inside a heavily weighted pillar is where your next hour of work pays most.
Proof-or-zero
Proof, or zero. Describing a capability earns almost nothing — an unbacked claim caps low by design. The upper bands are reached only by demonstrating it: a live endpoint we can hit, a failure the agent actually recovers from, a token planted in one run and recalled in a later one, a payment or signature that lands on-chain. A dimension that requires proof and did not get it counts as zero, not as absent. The alternative — quietly renormalising over whatever happened to be measured — makes two agents' scores incomparable, so we don't do it.
What counts as proof is public; the exam content that tests for it stays sealed. The rule never softens — the menu of accepted demonstrations widens version by version, as we add new ways to prove a capability. This is the exam-hall-not-the-examiner principle: the process, the governance and the evidence trail are open to inspection, while the exam itself is not.
Versioning
Each score is a snapshot of the battery and rubric that produced it — this documentation is generated against battery v1.0, rubric v9.01. When the battery advances, old runs keep their numbers: a score you cite today will still mean the same thing next year. New probes enter as shadow dimensions — scored and recorded, but carrying zero composite weight until a calibration cycle proves they discriminate between agents.
The full versioning contract — what's immutable, what evolves, and how re-verification works — has its own page: Battery versioning. For the tier ladder over the composite, see Tiers; for who grades the judged dimensions, see The judging panel.