No hype. Just what works in production.
The long-form home of @dataeng.ai. Carousels and reels compress ideas; this is where they get the full breakdown, the code, and the evidence.
Research to Production
Research results translated into production defaults. The papers, the benchmarks, and what to actually change on Monday.
3 POSTS
PILLAR 02Failure Modes / 2 AM Pager
What breaks at 2 AM and why. Incident-shaped lessons from production AI systems.
4 POSTS
PILLAR 03Tool Verdicts
Verdicts on the tools touching your data. What changed, what it costs, who owns the toggle.
3 POSTS

One email made Copilot leak private data, and the victim never clicked anything.
EchoLeak (CVE-2025-32711) slipped an attacker's email into Copilot's context past the injection filter, past link redaction, and out through an allowlisted Teams URL. The controls that failed all live in the retrieval layer.

A reranker gained 14 points on Hit@10. The answers barely moved.
A MedCPT cross-encoder lifted Hit@10 by 14 points on 1,000 clinical QA pairs, and answer correctness moved under 4. The authors' own words: retrieval gains do not translate proportionally into answer gains.

Your router might be costing you more.
RouteLLM's paper promises 2x savings on single-turn benchmarks. We measured what happens in agent loops: switching models drops cache reads to zero, and the 1.25x write price gets paid again every turn.

Rerunning an LLM step changes the output. Land it once, key it, and stop paying twice.
Temperature 0 was never deterministic. Treat every LLM call like a vendor API pull: land each response once, keyed on sha256 of input, model, and prompt version, and let backfills read the stored rows.

Your pipeline samples 2%. Judge all of them.
MotherDuck ran the numbers on judging all 100,000 rows: 89% accuracy, 40 seconds, fifty cents. The blocker was never accuracy. It was price.

Airflow 3.3.2 shipped an upgrade trap, and the error message lies.
pandas 3 renames the DataFrame class path and XComs record the old one. Upgrade Airflow on every worker first, or the error message will send you chasing your config.

dbt turned its AI on by default. Check the toggle.
From September 1, Wizard and Copilot ship enabled and every future dbt AI feature auto-enables. The meter is running, and the pool is shared.

Your demo worked. Your AI product won't.
Your demo ran on ten hand-picked examples. Production is the long tail of real inputs. Four things break, and none of them are the model.

Your support bot is hallucinating. Blame the dumpster, not the model.
2:14 AM. The bot is inventing refund policies. The retriever is fine and the prompt is fine. The problem is the fourteen thousand tokens of dumpster you fed it.


