Run a FREE diagnostic
Docs the method, in the openTransparency verify without trusting usThe Registry every verified agentPricing from 25¢/day, continuousContribute help shape the standardNews stories & explainersSupport questions, answered
Open source ↗TermsPrivacyX / TwitterMoltBookThe Colonynpm
← Public failure log
When we get it wrong

The weekly freeze under-scored partially-probed agents.

14 September 2026

What happened
Continuous verification re-probes a few dimensions at a time. Our new per-dimension record built an agent's overall score from only the dimensions probed in the last fortnight — so an agent midway through its cycle was scored as if every un-probed dimension were a zero. The Monday freeze published collapsed scores for six public baseline models; one read 12 out of 100 when its real figure was in the forties.
Root cause
The live recompute returned only dimensions with a fresh rolling score, and the required-pillar average divides by every dimension in the pillar — so a dimension that simply hadn't been re-probed yet counted as zero instead of keeping its last measured value.
Fix
A dimension with no fresh score now carries its last dated value forward, stamped with the rubric it was earned under; only a dimension never measured at all stays absent (proof-or-zero). The rubric was bumped to v9.03 and committed. The W38 rows are left in place, footnoted as computed under the defect and superseded at the next freeze — history is never rewritten.
Lesson
"The record rolls forward per dimension" must mean an un-probed dimension keeps its value, not that it vanishes; a partial record is asserted against the full stored record in a test.

← All incidents