
When Your Eval Script Answers a Question Nobody Asked
A scripted multi-turn eval replays fixed user turns regardless of what the agent said. The moment the agent asks a question the case author did not anticipate, the script answers something else and the case fails for the wrong reason - rewarding agents that guess and punishing agents that ask.



