When a model can't agree with itself
I ran the same model on the same 10-K five times with identical settings and asked it the same question each time. On one field it gave me the same answer five times out of five. On another field in the same filing it agreed with itself about seven times in ten.
The field it never wavered on was the board roster in a proxy statement. Agreement with itself: 1.0, across three hundred paired runs. Executive compensation cells from the same document came back at 0.995. The field it could not hold steady was the drivers a company names for its change in results, the narrative part of Management's Discussion and Analysis. There the same model, same document, same settings, agreed with its own prior answer 0.71 of the time. A cheaper model on that same field came in at 0.36.
The useful thing here is not that one model is better. It is that the same model produced both numbers, in the same run, on the same document. Whatever is going wrong at 0.71 is not a property of the model. It is a property of the question.
A board roster has an answer. The names sit in a table and either you read them correctly or you did not. "What drove the change in results" has no such answer written down anywhere in the filing. A human analyst asked that question twice a week apart would give two different answers too, and neither would be wrong. The model is not failing to retrieve a fact. It is being asked to make a judgment that the document does not settle, and it makes a slightly different one each time.
This matters because of what it does to comparison. Across model families, agreement on that MD&A field was 0.19. Read on its own, that number says the models fundamentally disagree about what the filing means, which sounds like a serious finding about model quality. Set it next to the 0.71 floor and it says something much more mundane and much more useful, which is that a good part of the disagreement is any single model disagreeing with itself. You cannot interpret disagreement between two instruments without knowing how much each one disagrees with itself. That floor is the first thing a measurement setup has to establish and it is almost never reported.
There is one more thing worth noticing in those five runs. The cost of each one was nearly identical, varying by less than a tenth of a cent. The number that stayed perfectly stable across all five repetitions is the number that tells you nothing about whether the answer was right. Every unstable thing was downstream of it.
So before you compare two extraction models, or trust a number that says they agree, find out what each one does when you ask it the same thing twice. If a field cannot survive that, no amount of model shopping will fix it, because the problem is in how the field was specified.
Read what an instrument like that measures, or the full benchmark.