AI & Cyber
Evaluating RAG Without Fooling Yourself
March 11, 2026·9 min read
The first RAG system I shipped scored 94% on our internal eval set and was quietly useless for three months before anyone noticed. The eval set was the problem.
The harness I use now
Adversarial queries, distractor documents, gold answers written by domain experts, and — crucially — a separate "is this even a question we should be answering" classifier.
— A.P.W.