Joshua Cook Consulting
BenchmarkWritingBook a call

Writing

On production AI systems that read documents, and on knowing whether to trust what they say.

  • August 27, 2026

    A green check is a statement about the check

    Every serious failure we hit last month was a check that passed while answering a different question than the one being asked.

  • August 3, 2026

    When a model can't agree with itself

    Running the same model on the same filing five times gave five different answers on one field and identical answers on another. The gap between those two is the most useful number in the benchmark.

  • July 27, 2026

    Poor men can't afford cheap shoes

    Extraction is cheap now, which is exactly the trap. A model that returns valid, empty JSON looks like the best deal and quietly costs you everything downstream.

  • July 19, 2026

    What extracting data from a filing actually costs in 2026

    Nine language models over 2,280 real SEC filings, every call metered. The extraction is close to free, and what you actually pay for is not knowing whether to trust it.

  • July 18, 2026

    The problem we were already living

    Every time software reads documents and reaches a conclusion, somebody has to answer an uncomfortable question. How do you know the call is right?

Author, Docker for Data Science and Pro Agentic AI (Apress). Instructor, Caltech CTME. MS Computer Science, Georgia Tech.

Joshua Cook Consulting · Independent AI extraction measurement for pharma · Remote, US