The Proxy Always Passes Its Own Test
When there is no cheap way to check whether something is true, the temptation is to check whether it agrees with itself. That is a different question wearing the real answer's clothes.
August 11, 2026 ยท Quantum Nexus Ventures FZCO
- verification
- legal AI
- e-discovery
- AI governance
In 2009 the U.S. National Institute of Standards and Technology ran the TREC Legal Track, an evaluation of methods for finding responsive documents in litigation. Its interactive task built in an appeal. Once the participating teams had submitted their runs, a sample of documents was coded by first-pass reviewers to build the answer key, and any team that believed a document had been coded wrongly could appeal it to the Topic Authority: a senior attorney playing the part of the lawyer who holds ultimate responsibility for deciding what is and is not relevant in a real production. That judgment was final. There was no second appeal.
Teams appealed about 5 percent of the assessed documents, and they appealed where they expected to win, because every overturned call raised their own score. Of those appeals, 89 percent succeeded. On the documents anyone bothered to contest, the senior attorney disagreed with the first-pass reviewer nearly nine times out of ten.
Not a gray area. Mostly plain error.
Maura Grossman and Gordon Cormack, two of the researchers who built the field, went back and asked why. Their finding, published in Pace Law Review in 2012, was not that responsiveness is inherently a matter of opinion. Re-reading the overturned documents themselves, they judged roughly nine in ten of them to be clearly responsive or clearly non-responsive, with only about 5 percent genuinely arguable. The disagreements were not close calls at the edge of a hard concept. They were mostly plain human error.
They did not spare the authority either. In 4 to 8 percent of the documents they re-read, the one they judged plainly wrong was the Topic Authority, not the reviewer. Grossman had served as Topic Authority on one of the topics herself; re-examining ten of her own documents, she reversed three and called two more arguable. And the paper states outright that its numbers cannot be used to estimate how often either the reviewer or the authority is wrong across the collection as a whole, because the sample was built deliberately out of contested documents. It is a study of where disagreement comes from, not a measurement of how often anyone is right.
Where the cheap measurement becomes the training signal
None of this was a problem for the evaluation itself. The teams never saw the first-pass coding; what they had was up to ten hours with the Topic Authority. The problem starts one step later, in ordinary practice, when the cheap measurement stops being a yardstick and becomes the thing a system is fitted to. Where a classifier learns from first-pass review coding, because that is the only labelled data anyone can afford to produce at volume, it inherits whatever that coding got wrong. A system trained to agree with the reviewer learns to agree with the reviewer. It does not learn to be right, and its accuracy score, computed against the same coding it was trained on, cannot tell the two apart.
The same shape, in our own audit trail
We found that shape in our own work, and it did not take an outside audit to catch it. It took two people on a different part of the team asking an uncomfortable question out loud. One asked who actually produced the evidence a check was running against. The other asked what the weakest possible input was that would still leave our confirmation sitting on true. We tested both questions against the real system instead of arguing about them, and the answer was worse than either question implied. A record built to prove that a legal claim had been verified correctly would still come back confirmed even when the underlying text it was checked against had been invented by whoever produced the record in the first place. It was internally consistent. Every part of it agreed with every other part. It had simply never been asked to agree with anything outside itself.
Self-consistency and authenticity are not the same property, and a single reassuring flag will always let the first stand in for the second, because whoever is reading it has no way to tell which one they were handed. So we took the flag apart. What passed before as one word now has to say specifically what it checked and what it did not, and a record that has not been anchored to anything outside its own claims says so, instead of defaulting to fine.
The input built to fool it
The lesson underneath both stories is the same one, fourteen years apart, in systems that have nothing else in common. When there is no cheap way to check whether something is true, the temptation is to check whether it agrees with itself, or with whoever is already inclined to agree with it. That is not a shortcut to the real answer. It is a different question wearing the real answer's clothes. And the only way to catch it is to go looking, on purpose, for the input built specifically to pass it, before someone else does.
Sources: Grossman, M. R., & Cormack, G. V. (2012). Inconsistent Responsiveness Determination in Document Review: Difference of Opinion or Human Error? 32 Pace Law Review 267. Hedin, B., Tomlinson, S., Baron, J. R., & Oard, D. W., Overview of the TREC 2009 Legal Track.Sources: Pace Law Review 32:2 ยท TREC 2009 Legal Track overview
This is an opinion / thought-leadership piece. It is not legal or financial advice.