Your pipeline samples 2 percent of rows and calls it a quality gate. Then it ships the other 98 percent on faith.
That was the deal we all accepted: judging every row was too slow and too expensive, so we sampled and hoped. MotherDuck ran the numbers on judging all 100,000 of them, and the numbers say the deal expired.
The comparison: an LLM judge scored 88 percent in 32 minutes for $37.58. The in-warehouse judge scored 89 percent in 40 seconds for $0.50. Slightly better accuracy, 48 times faster, 75 times cheaper.
The constraint was never accuracy
That is the sentence that matters. For years the objection to exhaustive judging was quality: a cheap judge would be a bad judge, so we sampled with the good stuff. The numbers say the cheap judge is as good as the expensive one. The blocker was never accuracy. It was price.
When a check costs fifty cents and forty seconds, it stops being a luxury and starts being a default. The question changes from “can we afford to judge everything” to “why are we still sampling.”
The judge moved into the warehouse
The trick is placement. prompt_jev() runs inside the warehouse, where the data already lives. No export, no API calls per row, no waiting on a model endpoint for 32 minutes. The data does not move. The judgment happens where the rows are.
This is the research-to-production pattern this pillar exists for. The research result (a decision model can judge at warehouse speed) becomes a production default (judge everything, quarantine in seconds). The architecture insight is boring and decisive: move the compute to the data, not the data to the compute.
What 100 percent judging buys you
Quarantines in 40 seconds, not 32 minutes. That is not a performance stat. It is an SLO stat. Forty seconds fits inside a five-minute freshness SLO with room to spare. Thirty-two minutes fails it outright, which means the LLM-judge pipeline cannot be a real gate. It is a report you read tomorrow.
Exhaustive judging turns data quality from an audit into a control. The bad rows get quarantined before the downstream consumer sees them, every run, not just on the sampled runs.
Provenance, stated plainly
Those are partner-run numbers, not ours, from the MotherDuck blog dated 2026-09-21, and the SQL is published so you can rerun it on your own labels. The caveats: 4-class news classification is not your schema, and there is no production case study yet. Rerun the published SQL on your labels before you quote these numbers to your team.
The point is not which model wins. The point is the blocker moved. When judging everything costs fifty cents, sampling is a habit, not a strategy.
This post started as an Instagram post →
Numbers above trace to these sources. If one moved, tell us and we fix it.


Talk it through
Argue with us on Instagram.