The problem we were already living
We built SchoolQualityReview, software that reads a school's documents and tells you how the school is doing. Strategic plans, board minutes, audits, the whole public record. It reads them and it makes a call.
Then a question shows up that has no comfortable answer. How do we know the call is right?
Not right in the sense of plausible. Not right in the sense of a school leader nodding along. Right in the sense of correct.
School quality has no answer key. Nobody has written down what this school actually deserves. There is no file somewhere that says this one is a 73 and that one is a 51, against which we could check our work. The whole reason SQR is useful is that this judgment is hard, contested, and expensive, which is exactly why nobody has already done it for every school and published the results.
So we could not check our answers. We could only check them against our own judgment, or against another model's judgment, which is a different way of saying the same thing.
That is not a school problem. That is what happens every time software reads documents and reaches a conclusion. A model reads a clinical note and decides a drug caused a reaction. A model reads a filing and extracts an obligation. A model reads a contract and flags a risk. In almost none of those cases does anybody have the answer key, because if the answers were already written down you would not need the model.
The field has an answer for this, and the answer is to use a proxy. Have another model grade the first one. Score how confident it sounds. Check whether the format is right. Measure agreement between runs. All reasonable. All things we did.
And then the uncomfortable question underneath: how do you know the proxy is right?
You are now measuring your measurement with another measurement, and at some point the tower has to rest on something. In a research setting you can stop the regress by buying labels, running the real answer, and checking. In production you cannot, because the reason you built the thing was that the real answer does not exist.
This is where I think the field has quietly gotten away with something. Extraction got so much better so quickly that "obviously better than what we had" started standing in for "good enough to trust." Those are not the same claim. The first one is easy to demonstrate. The second one requires knowing whether your proxy is any good, and almost nobody checks, because checking is hard and the improvement feels self-evident.
We did not come to this from the outside. We came to it because we had built a platform that made judgments nobody could grade, and we wanted to know whether we were right. That question turned out to be much harder and much more interesting than the extraction itself, and it turned out that the tools for answering it barely exist.
So we went and did the work. We took a benchmark that does have an answer key, we sealed it away where we could not see it, and we developed against nothing but proxies, exactly the way you have to develop in production. Then we opened the envelope and compared what the proxies said to what was true.
I will write about what we found when the paper is public. For now the part worth saying is the part that got us there.
If you are running extraction in production, you are steering with an instrument you have almost certainly never validated. That is not a criticism. There is no standard way to validate it. That is the actual gap, and it is the one worth caring about, because the instrument you cannot check is the one making your decisions.