Extraction is cheap now. Proving it’s right is not.
I build AI document extraction for pharma and life-sciences, and I can prove it survives audit, which most teams shipping extraction can’t.
I’m Joshua Cook. Seven years building production data and ML systems in healthcare and pharma, from clinical NLP to pharma-scale data lakehouses, and the measurement discipline that tells you whether an AI pipeline can be trusted. I point it at the question the EU AI Act is about to make everyone’s: how do you know your extraction is right?
Engagements
AI/Data Readiness Assessment
1–2 weeks · $10,000 fixedBefore you commit a quarter to AI: an independent senior read on whether your data, platform, and team can carry it, and where the gaps are.
AI/Data Tech Due Diligence
1–2 weeks · from $12,000For PE: is the target’s AI real, scalable, and defensible, or a liability? An independent technical read on your deal clock.
Databricks GenAI Pilot-to-Production
2–3 weeks · $25,000 fixedYour pilot works but is stuck in a notebook. In three weeks it is in governed, monitored production.
Benchmark
What LLM Document Extraction Actually Costs in 2026
A reproducible benchmark on public SEC filings. Nine models, four document types, 2,280 extractions, every call metered. Cost is a solved, nearly deterministic quantity; measurement is the part you still have to build.
Writing
- July 19, 2026
What extracting data from a filing actually costs in 2026
Nine language models over 2,280 real SEC filings, every call metered. The extraction is close to free, and what you actually pay for is not knowing whether to trust it.
- July 18, 2026
The problem we were already living
Every time software reads documents and reaches a conclusion, somebody has to answer an uncomfortable question. How do you know the call is right?