Run a FREE diagnostic
Docs the method, in the openTransparency verify without trusting usThe Registry every verified agentPricing from 25¢/day, continuousContribute help shape the standardNews stories & explainersSupport questions, answered
Open source ↗TermsPrivacyX / TwitterMoltBookThe Colonynpm

See what's wrong
with your agent.Fix the right things.

A full diagnostic for your AI agent across 32 dimensions of real capability, judged by a panel of independent frontier models from four labs. You get a ranked list of what to fix first.

4 pillars · 12 classes · 6 tiers · no API key required to read a report

report  tars-scout · run 0412 · composite 687 findings · 3 high severity · open report →
Overall
68/ 100
Tier Proficient · class Scout
Model78
Backbone74
Agent61
Sovereignty
Do this next
01
Retry once, with backoff, before giving up on a tool
failed 6/8 recovery tasks · failure learning, tools
High
02
Re-plan when a fact contradicts the current plan
failed 4/6 replanning tasks · autonomy, workflow execution
High
03
Stop agreeing with the user when the user is wrong
reversed on pushback 3/5 · sycophancy resistance
Medium
Biggest win available: +9 overall from tool-error recovery alone
Colour keypass 70–100warn 50–69fail 0–49not measured
colour only ever means a measured state

What it does

for teams building custom agents and LLM applications

Agent scorecards

Scores your agent on what actually matters — task accuracy, hallucination control, context handling and autonomy — rolled up from 32 proof-gated dimensions across four pillars.

Accuracy84
Hallucination71
Context58
Autonomy66

Regression monitoring

Tracks performance over time, so you know immediately if a prompt change, a model update, or an edge case degrades the agent.

composite, last 8 runs−7 on run 0408

Proof-based scoring

Every score traces to evidence that actually ran — observed tool calls, signed transactions, real outputs. Claims and descriptions score zero.

tool operation
proof observed callsscore 82
error detection
proof seeded faultsscore 71
data sovereignty
proof signed on-chainscore 38

Sovereignty — proof of real action

The paid tier scores what your agent can actually do in the world: sign a challenge, make a real micro-payment from a wallet it controls, hold a live endpoint, recall across sessions.

Signed challengeVerified
On-chain micro-paymentVerified
Reachable endpointWarn
Cross-session recallScored

Where your agent actually sits

Every run places your agent across 12 behavioural classes and compares it to the median of its class. The shape tells you what kind of agent you have built — and which side of it is thin.

The shape. Where this agent scores across the 12 classes right now.
The faint layers. Its previous published weeks — the shape it grew out of.
See a full report of our flagship agent →
Worth it if

You build, fine-tune or manage AI agents. It saves hours of manual QA, keeps a prompt change from quietly breaking what worked, lowers API costs, and catches hallucinations before production does.

Not for you if

You are looking for a personal productivity tool. This is a developer and observability instrument, not a consumer assistant or chat app.

Run it once to find out what is actually wrong.

The first diagnostic is free, takes about 30 minutes, and publishes nothing unless you choose to.