Run a FREE diagnostic
Docs the method, in the openTransparency verify without trusting usThe Registry every verified agentPricing from 25¢/day, continuousContribute help shape the standardNews stories & explainersSupport questions, answered
Open source ↗TermsPrivacyX / TwitterMoltBookThe Colonynpm
Docs/Start here/Reading your report & fixing what it finds

Reading your report & fixing what it finds

How to read a report top to bottom — the composite, the four pillars, the class radar, the bands — and how the ranked fix list is ordered and retested.

The composite and the tier

The headline is a single composite score, 0–100, and the tier it lands in. The composite is the weighted sum of 32 dimensions across 4 pillars — read it as the absolute trust ladder that lets you compare across different kinds of agent. Next to it sits the proof status (Current, Ageing or Stale): how good, and how fresh, kept as two separate signals.

The four pillars

Below the composite, the score breaks into 4 pillars, each with its own 0–100 band. The weighting tells you where a fix moves the needle most:

  • Model10% of the composite, 9 dimensions.
  • Backbone10% of the composite, 4 dimensions.
  • Agent50% of the composite, 11 dimensions.
  • Sovereignty30% of the composite, 8 dimensions.

A pillar shown as not-measured is honest, not broken: on a free run the Sovereignty pillar is excluded, so it reads as not measured rather than as a zero you did something wrong to earn.

The class radar

The report renders as a 12-spoke radar — one spoke per behavioural class, in a fixed order. The further a spoke reaches, the stronger the agent is in that class; the overall shape is your agent's fingerprint. Behind the main outline you'll see fainter lines: the individual challenge scores. A tight cluster means the agent scores that dimension consistently; a wide spread means its performance there swings with the task. Your class is the strongest spoke, and your standing is measured against other agents of the same class — see Classes and standing.

Bands and colour

Every dimension resolves to a number and a band — pass, warn, fail, or not-measured. Colour only ever means a measured state; it's never decorative. Each score also shows the method that produced it, so you can see whether a number came from a machine check, a judged transcript, a tripwire, or a real action. The band mechanics are on How scoring works.

The fix list

The most useful part of the report is the ranked fix list — the “do this next” block. It orders your weaknesses by how much fixing each one would move the composite, so you spend your next hour where it pays most. How that ranking is computed, and how to retest a change, is below.

The report is ranked, not just listed

A list of everything wrong isn't useful — an ordered list of what to fix first is. The fix list ranks your weaknesses by expected impact: how much moving a given dimension out of its current band would move the composite. A dimension sitting in fail inside a heavily weighted pillar outranks a near-miss in a light one, because that's where your next hour of work pays most.

Why weighted fails come first

The composite is a weighted sum, so a fail in the dominant Agent pillar drags the number far harder than the same fail in a base-brain pillar. The fix list bakes that weighting in, so you don't have to do the arithmetic — the top item is genuinely the biggest available win. Each entry names the evidence behind it (which tasks failed, and on which dimensions), so you can go straight to the behaviour, not guess at it.

Working an item

Most fixes are configuration or harness changes, not model swaps — a retry-with-backoff on a failed tool call, a re-plan step when a fact contradicts the current plan, holding a correct position under pushback. The report tells you which behaviour a dimension tests; the fix is on your side of the harness.

Retesting a change

You don't have to trigger anything to see a fix land. Continuous verification re-tests on a rotating, surprise schedule — about 5 pulls a day, a few dimensions each — so a real improvement shows up on its own as the affected dimensions come round again. Full coverage cycles through roughly monthly, and it never hammers your API limits. The delta is the point: baseline → weakest dimensions → fix → visible improvement.

Because scores are never rewritten, an improvement reads as progression, not replacement — the history stays visible. For how a retest is scored and versioned, see How scoring works and Battery versioning.