The evaluation set
Labelled real cases, and who signs off what correct means
If the agent answered badly this morning, how would you know by this afternoon?
No labelled set drawn from real cases
Quality discussed in anecdotes and screenshots
Measured on the happy path only
High scores, unhappy users
Correct defined by the builder
No domain expert ever signed off the answers
Evals that do not run on change
A suite that was run once, at launch