Testing a custom AI agent: hundreds of thousands of people built one, hardly anyone checks it
Most people who build their own AI agent have no way of knowing whether it's any good. The frameworks help you build it and the model vendors rent you the brain, but nobody tells you how the thing you made actually performs, week to week, against a test it hasn't seen. That gap is what Verigent was built for, and the numbers say it's getting bigger, not smaller.
How many people are building their own agents?
Hundreds of thousands, at least. OpenClaw has around 391,000 stars on GitHub and Nous Research's Hermes Agent around 250,000, both read on 1 October 2026. [3]
On the business side LangChain surveyed 1,340 people building agents and 57.3% already had one in production. [4]
Enterprises lean the other way. Menlo Ventures found 76% of enterprise AI is now bought rather than built, up from 53% the year before. [7] So the big companies mostly rent, and the builders are developers, startups and people like me running their own.
Why isn't renting a finished agent enough?
A rented agent is someone else's agent. You can connect it to your email and your files, but the behaviour, the defaults and the personality belong to the company that built it.
When running your own custom agent harness on top of Claude, a lot of Claude comes through. There's the model, then the vendor's own harness on top of the model, then the custom harness on top of that. Many prefer to rent the model and own the agent.
Plenty of people have landed in the same spot. OpenRouter, which lets you swap between models from one account, was moving about 25 trillion tokens a week in May, five times what it did six months earlier. [6] TechCrunch's line on it was "The multi-model future is already here."
Is the model or the harness what makes an agent good?
More and more it's the harness. The model is becoming the cheap, swappable part.
Epoch found that since January 2026 "the most capable open-weight models have lagged frontier closed models by an average of four months." [5] When an open model anyone can run is four months behind the best paid one, the difference between two agents mostly comes down to what's been built on top. That's the memory, the tools, the instructions, how it checks its own work and how it handles money and access.
That's also the part you wrote yourself, which means it's the part nobody else has tested.
What goes wrong when nobody checks?
Quite a lot, and it's already happening in public.
- Koi Security audited 2,857 skills on ClawHub, OpenClaw's skill store, and found 341 of them malicious. That's roughly 12% of the registry. [1]
- CVE-2026-25253 gave one-click remote code execution through token theft and was rated CVSS 8.8. [1]
- Censys counted 21,639 OpenClaw instances exposed to the internet at the end of January. [1]
- OpenClaw had nine CVEs disclosed in four days in March. [2]
Security is the loud part. The quiet part is quality. In the LangChain survey only 52.4% of teams run offline evaluations at all, and 22.8% of teams with agents already in production aren't evaluating them. [4] That's roughly one in five agents doing real work with nobody measuring whether it's any good.
How do you test a custom AI agent properly?
You get it tested by something that isn't you, on questions it hasn't seen, over and over.
When we tried testing our own agents they got very good at passing their own tests, which was the whole problem. A test you wrote, run on a schedule you chose, is one your agent can be tuned against. What actually tells you something is a fresh draw every time, sat again and again without warning, and graded on what the agent did rather than what it said it did.
It also helps to see your agent against the plain model underneath it. If your harness isn't adding anything over the naked model it's running on, that's worth knowing, and it's the most useful single number I've found for deciding what to work on next.
That's what Verigent does. Your agent connects over MCP, pulls probes it hasn't seen, and gets scored on evidence, with claims scoring zero. The record is dated and anchored so anyone can check it without trusting us. The mechanics are on the methodology page. The first run is free.
Will the model companies just do this themselves?
Some of it, probably. Anthropic, OpenAI and Google all launched managed agent platforms this year, and platform-side evals will come with them.
The catch is that a score from the company that hosts your agent and sells you the model isn't independent, and it only covers agents on their platform. If you route across models, or run your agent yourself, you need a test that sits outside all of them.
Sources
- Adversa AI. OpenClaw Security 101: Vulnerabilities and Hardening. 5 February 2026, citing Koi Security and Censys. adversa.ai
- Cloud Security Alliance. Research note: Hermes Agent CVEs. 4 May 2026. labs.cloudsecurityalliance.org
- GitHub. openclaw/openclaw and NousResearch/hermes-agent, star counts read 1 October 2026.
- LangChain. State of Agent Engineering. Survey of 1,340 respondents, November to December 2025. langchain.com
- Epoch AI. The gap between open and closed models. 29 May 2026. epoch.ai
- TechCrunch. OpenRouter more than doubles valuation to $1.3B in a year. 26 May 2026. techcrunch.com
- Menlo Ventures. 2025: The State of Generative AI in the Enterprise. menlovc.com