AI & Cyber

Evaluating RAG Without Fooling Yourself

March 11, 2026·9 min read

The first RAG system I shipped scored 94% on our internal eval set and was quietly useless for three months before anyone noticed. The eval set was the problem.

The harness I use now

Adversarial queries, distractor documents, gold answers written by domain experts, and — crucially — a separate "is this even a question we should be answering" classifier.

— A.P.W.

The Dispatch

New essays, straight to your inbox.

Occasional letters on network defense, AI risk, governance, and the life around the work. No spam, unsubscribe any time.