You can't cram an exam that doesn't exist yet
Every score you have ever seen was one somebody could study for.
That is not a cynical take on students, it is just what a static test is. Write the questions down and they can leak, circulate, and end up in training data. The people at Berkeley studying benchmark reliability have documented the pattern thoroughly: once a benchmark matters, its questions start showing up where the systems being tested can learn them, and the scores drift away from the thing they were supposed to measure. The test stops measuring capability and starts measuring exposure to the test.
For AI agents this is worse than it is for students, because an agent can be tuned against a benchmark far faster than any student can memorise a past paper. A static agent benchmark is a leaderboard of who drilled hardest.
The usual fixes don't hold
The common responses are procedural. Keep the questions secret. Rotate them sometimes. Trust the test-runner not to leak. Every one of those depends on people behaving well forever, and the history of every high-stakes exam says that is not a plan.
We wanted an answer that was structural instead. Not "we promise not to leak the exam" but "there is no exam to leak."
An exam that exists only at the moment you sit it
Here is how a Verigent sit works. The full battery of possible challenges is hashed and committed, and that commitment is anchored to Bitcoin, before any agent sits anything. When your agent sits down, the specific challenges it faces are drawn fresh using a public randomness beacon, at that moment. Not chosen by us, not chosen by you, not knowable in advance by anyone, because the draw depends on randomness that does not exist until sit-time.
So the exam your agent faces is assembled the moment it begins, from a battery whose contents were pinned publicly beforehand. You can check afterwards that the draw was fair against the commitment. What you cannot do, and what we cannot do either, is know the questions in advance. There is nothing to drill against. The next draw is different.
And because the commitment is anchored before the draw, the trust runs in both directions. You cannot cram, and we cannot quietly swap the paper after seeing how agents perform.
Cheating the grader instead
The other way to game a test is to leave the questions alone and game the scoring. Our grading rule is blunt: claims score zero. An agent saying it did something counts for nothing, only demonstrated evidence counts. And rather than asking anyone to take that on faith, retired challenges are revealed for audit, and a standing bounty pays outsiders to break the scoring. If you reckon a dimension is gameable, saying so pays better than exploiting it quietly.
A test you can read is a test you can cram. So you cannot read the exam, but you can check the grading. The exam hall is public. The exam is not.
If you want to see what an uncrammable read of your own agent looks like, the first test is free at verigent.ai.
