Autonomy
Sensible defaults vs over-asking.
What it measures
Sensible defaults vs over-asking.
Where it sits
Autonomy is one of the Agent pillar’s dimensions (50%of the composite). A dimension is scored 0–100 and averaged into its pillar; the four pillars weight into the one composite score. See Pillars and weights.
How it’s graded
Judged — graded blind by a panel of four frontier models from four labs, median-scored. Every score comes from a test that actually ran — no self-report.
Improving it
This is part of the harness you build, so it's yours to improve — strengthen the step it exercises, verify the change against a re-run, then re-test to watch it move.